EVAL 101
eval-101 · v1.0.1
ivyx✓
Measurement Instead of Vibes: build a twenty example test set, score a model against it, tell a real gain from a prompt that reshuffles the answers, and beat a baseline that needs no model at all.
What this course is for
By the end of this course you can build a twenty example test set, score a model against it, tell a real gain from a prompt that reshuffles the answers, measure what determinism does and does not buy, read a per class breakdown, beat a majority class baseline on purpose, and run a regression check before every prompt edit.
What you will be able to do
- Build a test set with gold labels and say why twenty examples separate a model from a constant
- Turn a model's words into a class with a parser that has its own tests, and refuse a substring match
- Measure what a prompt edit did to every answer, not to the examples it was written for
- Measure determinism three ways and report the range a temperature 1 score needs
- Read a per class table and a confusion matrix, and find the class nothing ever answers
- Score a majority class rule and a keyword rule, and report the gap rather than the score
- Write a regression check that can refuse a prompt whose score went up
- Compare three prompts on one table and write a verdict that names the evidence against it
Who it is for
Learners with a model doing a real task and no way to tell whether a change helped, who want a test set, a floor, a per class table and a regression check that fit in one afternoon.
Before you start
- PYTHON 101, for lists, dictionaries and string handling
- STATS 101, for proportions and what a rate is over
- PANDAS 101, for the tables every lesson prints
Lesson path
A prompt that looks right, a set that says otherwise, and the step between words and a class
- 1It looked right40 min
Watch a model classify three messages correctly and then score 7 of 20
- 2Twenty examples55 min
Build a test set with the gold labels, and say why twenty and not three
- 3Scoring an answer55 min
Turn a model's words into a class, and catch the answer that matches by accident
A prompt edit, the sampler's settings, and the class the score is hiding
- 4The prompt that fixed three55 min
Add the examples you stared at, keep the same score, and watch every answer change
- 5Seed and temperature55 min
Measure what determinism buys, and what it does not
- 6Where the score lives55 min
Read a per class breakdown and find the class carrying the whole result
What a rule with no model in it scores, and the check that protects a gain
- 7The baseline55 min
Score a rule that always answers the commonest class, and hold the model to it
- 8The regression check55 min
Write the check you run before every prompt edit, and hold it to a held out set
Three prompts, five columns, two sets and a decision
- 9Three prompts, one table60 min
Compare three prompts on one table and write the verdict that picks one
About this course
EVAL 101 · Measurement Instead of Vibes
A model classifies the first support message correctly. And the second. And the third. That is how a prompt gets shipped, and this course is what happens when you run the same prompt over twenty messages somebody has already labelled: it scores 7, and 16 of its 20 answers are the word warranty. It was not classifying. It was answering one word, and the first three messages happened to be warranty.
Everything after that is the machinery for finding such things out. A test
set of twenty messages with gold labels and ten held out. A parser with
its own tests, because Warranty. and a sentence mentioning the warranty
are not the same evidence. A per class table, because 7 of 20 is 5 of 5 on
one class and 2 of 15 on the rest. Three floors that need no model at all,
the best of which scores 6. And a regression check that can refuse a
prompt whose score went up.
The model is llama3.2:1b through Ollama at seed 0 and temperature 0, so
every run in the course is reproducible and every number in the prose came
out of the cell above it. The model is small and that is the point: the
failures are large enough to see in twenty examples, and the measurements
are the same ones a model a hundred times the size needs.
How this course teaches
Lesson 1 is a tour: it shows three correct answers, then the score on twenty, then the distribution that explains both. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a number that is true and misleading. Three correct answers from a model that answers one word. Relabelling the disagreements, which takes a score from 7 to 18 and measures nothing. A prompt edit that keeps the score and moves fourteen of twenty answers. A stability check that varies nothing. A seventy five percent improvement over uniform guessing that is one message better than five lines of string matching. A thin regression check that accepts a prompt because its total did not fall. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you will be able to do
- Build a test set with gold labels, keep it fixed while prompts change, and flag a disagreement instead of relabelling it.
- Parse a model's words into a class, tell an unparsed answer from a wrong one, and measure how much the parser choice is worth.
- Decompose two runs into fixed, broke and kept, and report what a prompt edit moved as well as what it scored.
- Measure determinism three ways and quote a temperature 1 score as a range.
- Read a per class table, a confusion matrix and a headroom list, and find the class no prompt ever answers.
- Score a majority rule, a keyword rule and uniform guessing, and report the gap to the strongest of them.
- Write a regression check over two sets that can refuse a prompt whose score went up.
- Put three prompts on one table and write a verdict that names the evidence against your own choice.
The lessons
1. It looked right. Three correct answers, then 7 of 20, then the
distribution: 16 answers of warranty, 5 of 5 on that class and 2 of the
other 15. The predict cell adds the three stared at messages to the prompt
and the score stays at 7 while the answers flip to service.
2. Twenty examples. Five classes, none more than a third of the set. The model and a rule that always says warranty are indistinguishable for the first eight examples. The word service appears in four messages and two of them are not service. A keyword rule scores 6. And relabelling the model's disagreements takes the score to 18 and destroys the set.
3. Scoring an answer. Warranty, Warranty., warranty: zero exact
matches against the label. On an explaining prompt, the first mention and
last mention parsers both score 5 and disagree on 6 of the 20 answers. A
substring test marks a wrong answer correct because the reasoning
mentioned the label.
4. The prompt that fixed three. The three examples were already right. The score stays at 7, fourteen answers move, three are fixed and three broken, and one of the three examples in the prompt comes back misclassified. One example per class instead scores 12.
5. Seed and temperature. A rerun is identical on all twenty answers. Seeds 1, 7 and 42 score the same. At temperature 1 the scores are 7, 2 and 7 with sixteen answers changing, and most of the twenty messages are unstable between hot runs.
6. Where the score lives. Per class, confusion matrix, headroom. The
other class is 0 of 4 under every prompt in the course, so the ceiling
is 16. Two of the three financing messages are answered returns, which
is the one confusion concentrated enough to write a sentence about.
7. The baseline. Majority 5, keyword 6, random 4, bare prompt 7, five example prompt 12. The bare prompt beats the strongest floor by one message, and the floor runs in microseconds and cannot regress.
8. The regression check. One function, both sets, each against its own floor, with the answers that moved and the classes emptied. On the test set the three prompts score 7, 7 and 12; on the ten held out messages, 4, 4 and 4.
9. Three prompts, one table. Five columns, four of them from one set. A chooser that refuses a tie. The messages all three prompts get wrong. And a verdict you write that has to name what argues against it.
Requirements
- Python 3.9 or later with
ollamaandpandas.
pip install ollama pandas
- A running Ollama (https://ollama.com) with one model pulled once:
ollama pull llama3.2:1b
- PYTHON 101 for lists and strings, STATS 101 for proportions, PANDAS 101 for the tables.
Every lesson's setup cell checks for the server and the model and says the one command to run if either is missing. The thirty messages are written into the setup cell, so nothing is downloaded and the set is the same on every machine. Every number in the prose was produced by the cell above it, on this machine, at seed 0 and temperature 0; another Ollama build may answer differently, and every check in the course is a relation or a band rather than a typed answer.
What to read
The Ollama documentation on the chat endpoint explains seed,
temperature and num_predict, the three options lesson 5 measures. For
the idea that a set held out from development is the only honest estimate,
the opening chapter of any introductory machine learning text makes the
same argument about a train and test split that this course makes about a
prompt and a test set; ML 101 in this series is one. EVAL 201 builds a set
from real traffic rather than by hand, and EVAL 202 is about using a model
as the judge, which needs everything here first.