EVAL 201
eval-201 · v1.0.0
ivyx✓
Building Eval Sets from Real Traffic: mine the hard cases out of a query log, write a labelling guide, measure two labellers above chance, revise the guide where the disagreement is, and sample a set you can defend.
What this course is for
By the end of this course you can mine an eval set out of a query log instead of inventing one: read what the log knows and what it does not, find the hard cases by margin and by abandonment, write a labelling guide, measure two labellers against each other above chance, revise the guide where the disagreement actually is, and sample a set that keeps the cases that matter.
What you will be able to do
- Count a week of traffic properly, separating rows from queries and keeping the repetition as data
- Tell an empty query, a search that returned nothing and a retry apart, and say what each one means for a set
- Find the rows worth labelling by margin and by abandonment, and judge a selection against what it leaves behind
- Write a labelling guide as numbered rules and measure which rule decides what, including the rows it is silent on
- Measure two labellers with percent agreement and the agreement chance alone would give
- Compute Cohen's kappa by hand, guard it where it is undefined, and read what it says about a set rather than a pair
- Find the classes a disagreement lives in, revise the guide to name them, and measure the revision away from the rows it was written from
- Design an eval set in strata, measure it against two random baselines, and write down what it can and cannot answer
Who it is for
Learners who have a system answering real queries and a log of what it did, and want an eval set mined from that traffic with labels two people agree on and a written record of what the set is for.
Before you start
- RAG 101, for the retriever whose answers this log records and the corpus behind it
- STATS 101, for proportions, samples and what a rate is over
- PANDAS 101, for grouping, filtering and joining the log to the key
Lesson path
A week of traffic, what it holds, and the rows worth a person's minute
- 1What people actually asked40 min
Read a week of traffic and find the questions nobody would have invented
- 2Reading the log55 min
Count what the log holds, from duplicates and empty queries to rows that returned nothing and a heavy tail
- 3Mining the hard cases55 min
Find the rows worth labelling by margin and by abandonment, and measure how much the two overlap
The guide, the two people following it, and what their agreement is worth
- 4The labelling guide55 min
Write the guide as three rules and apply it to the rows that decide whether the rules are enough
- 5Two labellers55 min
Measure percent agreement and the agreement chance alone would give
- 6Cohen's kappa55 min
Compute agreement above chance by hand and read what it says about the set
The rule the disagreement asked for, and the set it makes possible
- 7Revising the guide55 min
Find the classes the disagreement lives in, name them in the guide, and measure kappa again
- 8From log to eval set55 min
Sample a set that keeps the hard cases instead of the popular ones
Thirty questions, nine of them disputed, and two sentences you write
- 9Thirty rows you can defend60 min
Build a thirty example set from the log and defend two of its rows in writing
About this course
EVAL 201 · Building Eval Sets from Real Traffic
The usual way to build an eval set is to write twenty questions you can already answer, in the words the documents use, and watch the system pass them. This course builds one out of a week of real searches instead: 374 rows over eight days, 53 distinct queries, each with the document the retriever returned, how strongly it won, whether anybody clicked, and two labellers' verdicts.
Nineteen of the 47 questions with a known answer get the wrong document first, and they are not questions anybody would have invented: do you take part exchange against a corpus that says trade-in, how long is the guarantee against one that says warranty. The two labellers agree on 80.5 percent of the rows, which is a healthy looking number until you count what chance alone gives, 50.1, and then notice that on the sixty rows where the retriever returned another edition of the right document they agree on three.
No model is involved. The retriever is RAG 101's, the arithmetic is pandas and one formula written by hand, and the two labellers are rules rather than people, which the course says in its first lesson and again in the lesson that measures them. What is real is the shape of the disagreement, and finding that shape is the work.
How this course teaches
Lesson 1 is a tour: it reads the week, finds the questions the system answers badly, and ends by asking what two labellers would agree on by tossing coins. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a number that is defensible and misleading. Eighty percent agreement is half chance. A selection of the rows the retriever was most confident about is wrong 15 percent of the time and the rows it leaves behind are wrong 78. Conditioning agreement on one labeller's positive rows gives 0.95 and a kappa of 0.01. A revision scored on the rows that motivated it goes from 0 to 0.85. A majority over a question's rows turns 33 agreements and one dispute into a clean verdict. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you will be able to do
- Count a week of traffic properly: rows against queries, a heavy tail, and repetition kept as data rather than deduplicated away.
- Tell an empty query, a search that returned nothing and a retry apart, and say what each one means for a set built from them.
- Find the rows worth labelling by margin and by abandonment, and judge any selection by the rows it leaves behind rather than by its own error rate.
- Write a labelling guide as numbered rules, measure which rule decides what, and count the rows it is silent on.
- Measure two labellers with percent agreement and with the agreement their own habits would produce by chance.
- Compute Cohen's kappa by hand, guard it where it is undefined, and read it as a statement about a set rather than about a pair of people.
- Find the classes a disagreement lives in, revise the guide to name them, and measure the revision on a class it was not written from.
- Design a set in strata, measure it against two random baselines, and write down what it can and cannot answer.
The lessons
1. What people actually asked. 374 rows, 53 queries, eight days. The nineteen queries that get the wrong document first. Two labellers agreeing on 80.5 percent, and the chance level of 0.501 underneath it.
2. Reading the log. Fourteen queries are half the week and eight were asked once. Three empty queries, three searches that returned nothing, eight retries. Six queries the key has no question for, 6.4 percent of the rows. And the deduplication that produces 53 clean rows and throws the week away.
3. Mining the hard cases. Margins from 0.034 to 15.243. At a bar of one point, 151 rows and 76.6 percent wrong; abandoned, 199 rows and 82.3 percent; their union 242 rows at 72.2 against 2.4 for what it leaves. The rows the retriever was surest of are wrong 15.2 percent of the time, which is a fifth of the rate outside them.
4. The labelling guide. Three rules, one of which decides 94 percent of the log and one of which never fires. Twenty four rows set aside as out of scope, and the version that folds them into no. Sixty rows where the document that came back is another edition of the right one or another document of the same kind, which the guide never names.
5. Two labellers. The crosstab: 179 both yes, 122 both no, 64 rows where B says yes and A says no, 9 the other way. Agreement 0.949 on the rows that are not near misses and 0.05 on the sixty that are. Conditioning on A's yes rows gives 0.952 and means nothing.
6. Cohen's kappa. 0.6090, from 0.8048 observed and 0.5008 chance. Undefined on the 24 rows everybody agreed about. 0.570 when those rows are dropped, which is a Landis and Koch band lower for the same judgements. And 0.01 on a subset with 0.95 agreement.
7. Revising the guide. Rule 4 and its note: a near miss is not a hit, and the relationship is recorded. Agreement 0.805 to 0.955, kappa 0.609 to 0.909, the untouched rows moving by 0.006, and the sixteen same kind rows the rule was not written from going from 0.0 to 1.0. Seventeen rows survive, seven of them slips on a right answer.
8. From log to eval set. Thirty rows at random are nineteen questions. Thirty queries at random cover 221 rows and are 47 percent wrong. Twelve busiest, twelve hardest, four rarest and every remaining unscored query is thirty questions, 238 rows and 60 percent wrong, and the last stratum asked for six and got two.
9. Thirty rows you can defend. Nine of the thirty questions are disputed by ten rows, and they are mostly the busiest ones, so the set covers 238 rows and can score about 100 of them today. A queue for a third person, a note per near miss, and two sentences you write.
Requirements
- Python 3.9 or later with
pandasandnumpy.
pip install pandas numpy
- RAG 101, for the retriever whose answers this log records; STATS 101 for proportions and samples; PANDAS 101 for grouping and joining.
No model, no network and no API key. Every lesson's setup cell rewrites the five files it needs: the query log, the relabelled verdicts, the mapping from query to question, the corpus index and RAG 101's key. The random baselines use a fixed seed, so the numbers in the prose are the numbers your notebook prints.
What to read
Cohen's 1960 paper, A Coefficient of Agreement for Nominal Scales, is four pages and the source of lesson 6's formula; the paragraph on why the raw percentage misleads is this course's lesson 5 in one page. Landis and Koch's 1977 bands are quoted as thresholds far more often than they were meant to be, which lesson 6 demonstrates by moving a labelling job between two of them without changing a judgement. Aroyo and Welty's Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation argues that disagreement between labellers is signal rather than noise, which is exactly what lesson 7 does with it.