COST 201
cost-201 · v1.0.1
ivyx✓
Unit Economics: price a feature per request from the tokens it really uses, measure what a cache is worth on real traffic, compare tiers including the one with no model in it, and get the thing into profit.
What this course is for
By the end of this course you can price a feature per request from the tokens it really uses, read the traffic shape that decides the bill, measure what a cache is worth on real queries, compare tiers that include not calling a model at all, and get a feature into profit or say why it cannot be.
What you will be able to do
- Read a model's own token counters and sweep the context size to see which half of the bill moves
- Keep a price table in one place with its date, and know which conclusions survive a different one
- Price a week of real traffic by its counts rather than by its question list
- Measure a cache hit rate by replaying traffic, and separate a warm up from a steady state
- Compare four tiers on accuracy and cost per right answer, including the tier with no model call
- Measure a reasoning model's real output and find that an output cap is not a cost control
- Turn costs into margins and break even prices, per model call and per request
- Choose a design under an accuracy floor and a margin floor, and name what the table cannot measure
Who it is for
Learners with a feature that calls a model and no idea what one answer costs, who want the tokens, the traffic, the cache and the margin measured rather than estimated.
Before you start
- RAG 101, for the retriever every tier in this course is built on
- EVAL 201, for the query log this course prices
- PANDAS 101, for the tables every lesson prints
Lesson path
One answer, its tokens, and the price table that turns them into money
- 1The feature, priced40 min
Put a token count and a price on one answer and find the free tier beating the paid one
- 2Tokens in and out55 min
Read the model's own counts, sweep the context size, and see which half of the bill moves
- 3The price table55 min
Turn tokens into money with an assumption you can change in one place
A week of real queries, the cache that removes most of them, and the tiers that answer the rest
- 4Cost per user55 min
Price a week of real queries and read the heavy tail that decides the total
- 5The cache55 min
Measure the hit rate on real traffic and price the request that never happens
- 6Tiering55 min
Compare four tiers on accuracy and cost, including the tier with no model in it
A reasoning model's real invoice, and every design against a selling price
- 7The reasoning invoice55 min
Measure a reasoning model's output tokens and find that asking it not to think does not stop it
- 8The margin table55 min
Put every design on one table with its accuracy, its cost per thousand requests and its margin
Every lever measured on one week, and a design chosen under two floors
- 9Into profit60 min
Choose a design that answers well enough and earns money, and name the lever that did it
About this course
COST 201 · Unit Economics
Northgate's search answers a question by retrieving three passages and asking a model for one sentence. Nobody had asked what one answer costs. It reads 335 tokens and writes 20, takes a third of a second, and under the price table this course assumes it costs a fifth of a cent. A week of the real traffic is eight cents.
Then the second measurement, which is the one this course is really about. Over the 24 questions with a known answer, the model gets 14 right. Returning the top passage with no model call at all gets 17. The free tier is more accurate than the paid one on this corpus, and every optimisation in the seven lessons after the tour is a way of paying less for something that was already losing to one line of code.
The models are local, llama3.2:1b and qwen3:4b through Ollama at seed
0 and temperature 0, and every token count is the model's own
prompt_eval_count and eval_count. The price table is the course's own
assumption, stated as one and kept in a single dictionary, because a
published price list goes stale faster than a course does.
How this course teaches
Lesson 1 is a tour: it prices one request, prices the week, and then scores the tier with no model in it. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a number that is arithmetically right and about the wrong thing. A character based token estimate. A price written into an expression that will be copied four times. A week priced from its question list rather than its request log, which understates it sevenfold. A cache hit rate measured while the cache was filling. A margin computed per model call and compared with a price charged per request. An output cap that halves the bill by removing the answers. And a tier chosen by sorting on cost with the accuracy column never consulted. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you will be able to do
- Read a model's own counters, sweep the context size, and say which half of the bill each design change moves.
- Keep a price table in one place with its date beside it, and test which conclusions survive a different table.
- Price a week of real traffic by its counts, and give a cost per user with its assumption attached.
- Measure a cache hit rate by replaying the log, price a freshness policy, and separate the warm up from the steady state.
- Compare tiers on accuracy, latency and cost per right answer, and route part of the traffic to a cheaper one.
- Measure what a reasoning model actually writes, and why an output cap is not a cost control on one.
- Turn costs into margins and break even prices, per model call and per request, and show what volume does to a negative margin.
- Choose a design under two floors and name the column the table cannot hold.
The lessons
1. The feature, priced. 366 tokens in, 22 out, a third of a second, a fifth of a cent. Four fifths of it is the context. 14 of 24 right, against 17 for the retrieved passage with no model call.
2. Tokens in and out. 117, 335, 548 and 1,000 input tokens at one, three, five and eight passages; 16 to 21 output tokens at all of them; 15, 14, 14, 15 right. Six times the money across the sweep and nothing to show for it. A character estimate measured against the counter.
3. The price table. Two rows, four numbers, output priced four times input. A uniform price rise reorders nothing; only the ratio can. A budget in requests per hundred dollars, and a price written into an expression refused.
4. Cost per user. 374 requests, 53 distinct queries, seven requests per query, so a bill built from the question list is a seventh of the real one. The top five queries are a quarter of the week and the queries asked once are a twentieth.
5. The cache. 85.8 percent hits, 14 percent of the bill left, a normalised key worth nothing on this traffic, a daily expiry worth a quarter of a permanent cache, and a warranty question answered wrongly 34 times for the price of one.
6. Tiering. 17 right at nothing, 15 at nine cents per thousand, 14 at twenty one, 14 at two dollars seven. Routing the head saves a quarter before a cache and five requests after it, because levers multiply.
7. The reasoning invoice. 365 output tokens against 20, seven seconds against a quarter, one more right answer in six. The thinking flag moves the reasoning into the content field rather than suppressing it, and a 200 token cap is paid in full and returns nothing.
8. The margin table. At fifty cents per thousand, three designs are profitable per model call and four per request behind a cache; the reasoning design is not, and at a million requests a month it loses eight thousand dollars. Break even is the cost under another name.
9. Into profit. Eight designs, two levers stacked to six percent of the baseline, and a verdict you write that has to name the objection the table cannot hold: the free design returns a paragraph where the model returns a sentence.
Requirements
- Python 3.9 or later with
ollama,pandasandnumpy.
pip install ollama pandas numpy
- A running Ollama (https://ollama.com) with two models pulled once:
ollama pull llama3.2:1b
ollama pull qwen3:4b
- RAG 101 for the retriever, EVAL 201 for the query log this course prices, PANDAS 101 for the tables.
Only lesson 7 needs the larger model, and it says so. Every lesson's setup cell checks for the server and the models it uses and names the pull command. Every number in the prose was produced by the cell above it on this machine; the price table and the selling price are assumptions the course states, and every figure derived from them moves when you change them.
What to read
The Ollama documentation on the chat endpoint explains prompt_eval_count
and eval_count, the two numbers every cost in this course is built on,
and the think option lesson 7 measures. Hosted providers publish their
price tables in dollars per million tokens; the durable advice is the one
this course follows, which is to keep yours in one place with a date on
it. RAG 101's lesson 6 is where the retrieved passage first turned out to
contain the answer, CTX 201 is why the model sometimes reads the right
passage and answers from the wrong sentence, and EVAL 201 is where the
query log priced here came from.