← All courses
IVYXSTUDIO · COURSE

SERVE 201

serve-201 · v1.0.0

ivyx✓

Serving and Scaling an AI Service: run traffic through a simulated service on a virtual clock, watch the latency curve go flat then vertical, learn what a queue buys and what an unbounded one costs, size the concurrency limit against the invoice, then run a real service once and judge the numbers.

intermediate485 min9 lessonsen
#serving#scaling#latency#queueing#load-shedding#platform#ai-series#intermediate

What this course is for

By the end of this course you can say what a queue is actually buying, pick a concurrency limit and an operating utilisation from a measurement rather than a round number, tell a service shedding load on purpose from one failing slowly, and judge which numbers from a simulation survive a real run.

What you will be able to do

  • Predict the latency at eighty per cent utilisation, measure it, and find the curve is not a line
  • Separate absorbing a burst from absorbing a rate and name which a queue cannot do
  • Choose a concurrency limit from the curve and defend it against the numbers either side
  • Run one overload under three policies and report answered in time rather than answered
  • Measure the load a timeout adds when the work carries on after the caller has left
  • Refuse early and on purpose with a refusal the caller can act on
  • Compare time to first token against total time and find the policy that trades one for the other
  • Put the price beside the curve, pick the operating point, and judge whether the simulation was honest

Who it is for

Learners who finished CONT 201 and COST 201 and now operate the service under load: sizing it, queueing it, and deciding how it fails when it cannot keep up.

Before you start

  • CONT 201, for the container the service runs in
  • COST 201, for the token counts this course turns into service times and prices

Lesson path

Load3 lessons

The utilisation curve, the queue, and the concurrency limit

  1. 1Twice the servers, the same wait40 min

    Predict the latency at eighty per cent utilisation, then measure it, and find the curve is not a line

  2. 2What a queue is buying55 min

    Separate absorbing a burst from absorbing a rate, and name which one a queue cannot do

  3. 3Picking a concurrency limit55 min

    Choose a number from the measured curve and defend it against both the round number above it and the one below

Failing well3 lessons

The unbounded queue, the missing cancel, and backpressure

  1. 4The unbounded queue55 min

    Run one overload under three policies and report answered in time rather than answered

  2. 5A timeout without a cancel55 min

    Measure the load a timeout adds when the work carries on after the caller has left

  3. 6Backpressure55 min

    Refuse early and on purpose, and show the caller a refusal they can act on

The shape of the work1 lesson

Streaming, and the two latencies it splits apart

  1. 7Streaming changes the question55 min

    Compare time to first token against total time and find the policy that improves one by making the other worse

Judgment2 lessons

The invoice beside the curve, and whether the simulation was honest

  1. 8Size it against the invoice55 min

    Put COST 201's price beside the latency curve and pick the operating point, naming what you are buying

  2. 9Was the simulation honest60 min

    Run the real service once, compare it with the simulation, and say which of the course's numbers you would still trust

About this course

SERVE 201 · Serving and Scaling an AI Service

The courses before this one built a service and ran it for one customer, then two. This one runs it under load and asks the question that decides whether it stays up: how many requests a second can it take before the wait becomes unusable, and what does adding capacity actually buy. The answer is not the one most people carry. Latency against load is flat and then vertical, so a service sized at a comfortable margin is one traffic spike from unusable, and near the flat part more servers change almost nothing.

The whole course is a simulation, on purpose. A real service's latency depends on the machine, the hour and the neighbours, so no two runs agree and no number is teachable. The simulator here runs on a virtual clock: it plays traffic through a service with a fixed number of servers and a queue, and because the clock is virtual and the arrivals are seeded, every number is the same on every machine. The arrival shape comes from the query log, the per-request service time from COST 201's token counts, and the price from COST 201's table. The last lesson runs a real service once beside the simulation and asks which of the course's numbers survive a real clock.

How this course teaches

Lesson 1 is a tour: it sweeps utilisation, measures the median and tail latency at each point, and finds the curve is flat then vertical rather than a line. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.

  • A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
  • Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
  • An exercise that is broken when you open it.
  • A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
  • A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.

No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.

The particular danger of this subject is a metric that reads well while the service fails. An unbounded queue answers a hundred per cent of requests and delivers almost none in time. A caller-side timeout with no server cancel looks like protection and wastes thousands of server-seconds on abandoned work. A cost dashboard always points at running the service hotter, straight into the vertical part of the latency curve. Every diagnose cell is one of those, and every cross check is the second computation that refuses it.

What you will be able to do

  • Read a latency-against-utilisation curve, find the knee, and know why more servers buy the least where the service is comfortable.
  • Tell a burst from a rate and say which one a queue can absorb and which it cannot.
  • Pick a concurrency limit off the curve and defend it against the numbers on either side.
  • Report answered in time rather than answered, and see why an unbounded queue is the worst policy under overload.
  • Measure the capacity a timeout without a cancel wastes, and wire the cancel that frees it.
  • Refuse early and legibly, with a retry-after the caller can act on.
  • Split a streaming service's latency into time to first token and total, and choose the policy against the one the caller feels.
  • Put the price beside the curve, pick the operating point at the knee, and judge which of a simulation's numbers survive a real run.

The lessons

1. Twice the servers, the same wait. The p99 is flat from 0.5 to 0.8 utilisation, 4.1 to 5.5 seconds, then vertical, 5.5 to 10 to 24 at 0.9 and 0.95. Doubling the servers at a comfortable load barely moves the wait, because there it is mostly service time, not queue.

2. What a queue is buying. A burst of 400 into an idle service drains in about two minutes, refusing nobody; a sustained rate at 1.3 times capacity refuses nobody either but the wait climbs from 15 seconds to over 400 and keeps going. A queue absorbs a burst, never a rate.

3. Picking a concurrency limit. At a fixed rate, three servers overload the service to a 208-second p99, four sit on the knee at 6 seconds, five have headroom at 4, and eight barely beat five for double the machines of four. The pick is five, defended against both.

4. The unbounded queue. Under a 50 per cent overload the unbounded queue answers 100 per cent and delivers 0.3 per cent in time; a bounded queue of two delivers 60, immediate refusal 51. Answered and answered-in-time invert the ranking.

5. A timeout without a cancel. The same overload burns 5,335 server-seconds finishing requests whose callers had left, delivering 12 in time; adding a cancel delivers 371, thirty times as many. A timeout is the caller's decision, a cancel is the server's.

6. Backpressure. The backpressure service refuses 1,473 early and answers 2,386 in time; the accept-all service misses 3,988 at the deadline. A refusal needs a retry-after the caller can act on, or the refused callers retry at once and undo the shedding.

7. Streaming changes the question. First-in-first-out gives a 0.76-second first token and a 2.13-second total; processor-sharing gives 0.05 and 2.32, a fifteenfold faster first token for a ten per cent slower total. The right policy depends on what the caller feels.

8. Size it against the invoice. Cost per request falls 14 per cent from 0.7 to 0.95 utilisation while the p99 rises fivefold. The operating point is the knee, where cost is near its floor and latency still flat, and the cost curve alone always points at the wall.

9. Was the simulation honest. A real BM25 service through a thread pool clears the same throughput on one thread and eight, because Python holds the interpreter lock, so the simulation's capacity-scaling assumption fails for CPU-bound work and holds for lock-releasing work. The shape results survive; the capacity results are conditional.

What you need

This course runs on Python with numpy and pandas. Install them into the kernel with:

%pip install numpy pandas

There is no model, no server and no network call. The simulator, the arrival sequence and the price table are all built in the first cell of every lesson, so each lesson runs on its own. The last lesson runs a real thread pool for a couple of seconds; its numbers vary from run to run and are never graded.

What to read outside this course

CONT 201 built the container the service runs in, and COST 201 measured the token counts this course turns into service times and prices, so both are the direct prerequisites. The classical name for the flat-then-vertical curve is queueing theory, and any treatment of the M/M/c queue derives the shape this course measures; the value of measuring it here is that the simulator lets you change the servers, the deadline and the queue policy and watch the curve move. The literature on tail latency, the p99 rather than the average, is where the choice to report the ninety-ninth percentile comes from.