← All courses
IVYXSTUDIO · COURSE

COST 201

cost-201 · v1.0.1

ivyx✓

Unit Economics: price a feature per request from the tokens it really uses, measure what a cache is worth on real traffic, compare tiers including the one with no model in it, and get the thing into profit.

intermediate485 min9 lessonsen
#cost#unit-economics#tokens#caching#tiering#ollama#ai-series#intermediate

What this course is for

By the end of this course you can price a feature per request from the tokens it really uses, read the traffic shape that decides the bill, measure what a cache is worth on real queries, compare tiers that include not calling a model at all, and get a feature into profit or say why it cannot be.

What you will be able to do

  • Read a model's own token counters and sweep the context size to see which half of the bill moves
  • Keep a price table in one place with its date, and know which conclusions survive a different one
  • Price a week of real traffic by its counts rather than by its question list
  • Measure a cache hit rate by replaying traffic, and separate a warm up from a steady state
  • Compare four tiers on accuracy and cost per right answer, including the tier with no model call
  • Measure a reasoning model's real output and find that an output cap is not a cost control
  • Turn costs into margins and break even prices, per model call and per request
  • Choose a design under an accuracy floor and a margin floor, and name what the table cannot measure

Who it is for

Learners with a feature that calls a model and no idea what one answer costs, who want the tokens, the traffic, the cache and the margin measured rather than estimated.

Before you start

  • RAG 101, for the retriever every tier in this course is built on
  • EVAL 201, for the query log this course prices
  • PANDAS 101, for the tables every lesson prints

Lesson path

What a request costs3 lessons

One answer, its tokens, and the price table that turns them into money

  1. 1The feature, priced40 min

    Put a token count and a price on one answer and find the free tier beating the paid one

  2. 2Tokens in and out55 min

    Read the model's own counts, sweep the context size, and see which half of the bill moves

  3. 3The price table55 min

    Turn tokens into money with an assumption you can change in one place

What the traffic does3 lessons

A week of real queries, the cache that removes most of them, and the tiers that answer the rest

  1. 4Cost per user55 min

    Price a week of real queries and read the heavy tail that decides the total

  2. 5The cache55 min

    Measure the hit rate on real traffic and price the request that never happens

  3. 6Tiering55 min

    Compare four tiers on accuracy and cost, including the tier with no model in it

What you are billed for2 lessons

A reasoning model's real invoice, and every design against a selling price

  1. 7The reasoning invoice55 min

    Measure a reasoning model's output tokens and find that asking it not to think does not stop it

  2. 8The margin table55 min

    Put every design on one table with its accuracy, its cost per thousand requests and its margin

Judgment1 lesson

Every lever measured on one week, and a design chosen under two floors

  1. 9Into profit60 min

    Choose a design that answers well enough and earns money, and name the lever that did it

About this course

COST 201 · Unit Economics

Northgate's search answers a question by retrieving three passages and asking a model for one sentence. Nobody had asked what one answer costs. It reads 335 tokens and writes 20, takes a third of a second, and under the price table this course assumes it costs a fifth of a cent. A week of the real traffic is eight cents.

Then the second measurement, which is the one this course is really about. Over the 24 questions with a known answer, the model gets 14 right. Returning the top passage with no model call at all gets 17. The free tier is more accurate than the paid one on this corpus, and every optimisation in the seven lessons after the tour is a way of paying less for something that was already losing to one line of code.

The models are local, llama3.2:1b and qwen3:4b through Ollama at seed 0 and temperature 0, and every token count is the model's own prompt_eval_count and eval_count. The price table is the course's own assumption, stated as one and kept in a single dictionary, because a published price list goes stale faster than a course does.

How this course teaches

Lesson 1 is a tour: it prices one request, prices the week, and then scores the tier with no model in it. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.

  • A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
  • Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
  • An exercise that is broken when you open it.
  • A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
  • A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.

No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.

The particular danger of this subject is a number that is arithmetically right and about the wrong thing. A character based token estimate. A price written into an expression that will be copied four times. A week priced from its question list rather than its request log, which understates it sevenfold. A cache hit rate measured while the cache was filling. A margin computed per model call and compared with a price charged per request. An output cap that halves the bill by removing the answers. And a tier chosen by sorting on cost with the accuracy column never consulted. Every diagnose cell is one of those, and every cross check is the second route that refuses it.

What you will be able to do

  • Read a model's own counters, sweep the context size, and say which half of the bill each design change moves.
  • Keep a price table in one place with its date beside it, and test which conclusions survive a different table.
  • Price a week of real traffic by its counts, and give a cost per user with its assumption attached.
  • Measure a cache hit rate by replaying the log, price a freshness policy, and separate the warm up from the steady state.
  • Compare tiers on accuracy, latency and cost per right answer, and route part of the traffic to a cheaper one.
  • Measure what a reasoning model actually writes, and why an output cap is not a cost control on one.
  • Turn costs into margins and break even prices, per model call and per request, and show what volume does to a negative margin.
  • Choose a design under two floors and name the column the table cannot hold.

The lessons

1. The feature, priced. 366 tokens in, 22 out, a third of a second, a fifth of a cent. Four fifths of it is the context. 14 of 24 right, against 17 for the retrieved passage with no model call.

2. Tokens in and out. 117, 335, 548 and 1,000 input tokens at one, three, five and eight passages; 16 to 21 output tokens at all of them; 15, 14, 14, 15 right. Six times the money across the sweep and nothing to show for it. A character estimate measured against the counter.

3. The price table. Two rows, four numbers, output priced four times input. A uniform price rise reorders nothing; only the ratio can. A budget in requests per hundred dollars, and a price written into an expression refused.

4. Cost per user. 374 requests, 53 distinct queries, seven requests per query, so a bill built from the question list is a seventh of the real one. The top five queries are a quarter of the week and the queries asked once are a twentieth.

5. The cache. 85.8 percent hits, 14 percent of the bill left, a normalised key worth nothing on this traffic, a daily expiry worth a quarter of a permanent cache, and a warranty question answered wrongly 34 times for the price of one.

6. Tiering. 17 right at nothing, 15 at nine cents per thousand, 14 at twenty one, 14 at two dollars seven. Routing the head saves a quarter before a cache and five requests after it, because levers multiply.

7. The reasoning invoice. 365 output tokens against 20, seven seconds against a quarter, one more right answer in six. The thinking flag moves the reasoning into the content field rather than suppressing it, and a 200 token cap is paid in full and returns nothing.

8. The margin table. At fifty cents per thousand, three designs are profitable per model call and four per request behind a cache; the reasoning design is not, and at a million requests a month it loses eight thousand dollars. Break even is the cost under another name.

9. Into profit. Eight designs, two levers stacked to six percent of the baseline, and a verdict you write that has to name the objection the table cannot hold: the free design returns a paragraph where the model returns a sentence.

Requirements

  • Python 3.9 or later with ollama, pandas and numpy.
pip install ollama pandas numpy
ollama pull llama3.2:1b
ollama pull qwen3:4b
  • RAG 101 for the retriever, EVAL 201 for the query log this course prices, PANDAS 101 for the tables.

Only lesson 7 needs the larger model, and it says so. Every lesson's setup cell checks for the server and the models it uses and names the pull command. Every number in the prose was produced by the cell above it on this machine; the price table and the selling price are assumptions the course states, and every figure derived from them moves when you change them.

What to read

The Ollama documentation on the chat endpoint explains prompt_eval_count and eval_count, the two numbers every cost in this course is built on, and the think option lesson 7 measures. Hosted providers publish their price tables in dollars per million tokens; the durable advice is the one this course follows, which is to keep yours in one place with a date on it. RAG 101's lesson 6 is where the retrieved passage first turned out to contain the answer, CTX 201 is why the model sometimes reads the right passage and answers from the wrong sentence, and EVAL 201 is where the query log priced here came from.