EVAL 202
eval-202 · v1.0.0
ivyx✓
LLM as Judge: hold a judge model to answers a person already scored, give it the facts, measure its length and position bias, translate its scale, and say in writing which uses it is fit for.
What this course is for
By the end of this course you can write a judge rubric and hold the judge to answers a person already scored, read five agreement numbers and say which one hides an offset, find the facts a judge needs in its prompt, measure length and position bias and undo what can be undone, translate the judge's scale into a person's, and say which uses a judge is fit for before you trust a number from it.
What you will be able to do
- Build a judged set with a person's scores and a second reading, and measure the floor a constant judge sets
- Split agreement into exact, within one, error, rank correlation and kappa, and say what each one cannot see
- Put the gold answer and then the document into the judge's prompt, and measure what the wrong answers and the system ranking do
- Pad every answer with content free text, count the scores that rose, and price the rubric sentence that stops it
- Judge every pair in both orders, count the flips, find each judge's lean, and hold the surviving verdicts to the person
- Translate the judge's scale through a calibration table, and choose a pass threshold and a flag bar on one half and hold them on the other
- Measure a rerun, three hot seeds and a majority vote, and the range a rubric's wording moves the number across
- Put ranking, flagging and grading on one table and write the verdict that names the uses the judge is fit for
Who it is for
Learners who have a model producing free text answers and want a second model to score them, and who need to know before trusting the score what it correlates with, what moves it, and which decisions it can carry.
Before you start
- EVAL 101, for the test set, the scorer, the baseline and the regression check this course builds on
- EVAL 201, for where human scores come from and how much two people agree
- PANDAS 101, for the tables every lesson prints
Lesson path
Two judges, a set a person scored, and five ways to say how much they agree
- 1A model grades a model40 min
Watch two judges score six answers a person already scored, and predict how the better one does over forty eight
- 2Forty eight answers and a rubric55 min
Build the judged set with its human scores, write the rubric, and find the judge that answers one number
- 3Agreement is five numbers55 min
Measure exact, within one, error, rank correlation and kappa, and find the one that hides an offset
The facts it was never shown, the padding it rewards, and the slot it prefers
- 4The judge without the facts55 min
Put the document in the prompt and watch the wrong answers fall and the systems change places
- 5Longer looks better55 min
Pad every answer with nothing and count the scores that rose, then measure what the fix costs
- 6Which one came first55 min
Judge every pair in both orders, count the flips, and find the lean each judge has
What the judge's four means, and what moves it
- 7What a four means55 min
Map the judge's scale onto the person's, pick a threshold on half the set and hold it on the other half
- 8The same verdict twice55 min
Measure a rerun, a hot seed and a majority vote, then change one clause of the rubric and watch the ranking move
Three uses, two judges, one table and a verdict
- 9Three uses, one judge60 min
Put ranking, flagging and grading on one table and write the verdict that says which the judge is fit for
About this course
EVAL 202 · LLM as Judge
A service desk answer is a sentence, and is this a good answer has no label to match it against. The usual move is to hand it to a second model with a rubric and take the score. This course holds that score to forty eight answers a person already scored, and the first thing it finds is that the judge gives two confidently wrong answers a 5, ranks the worse of two systems above the better, and marks its own answer key wrong nine times out of twenty four.
Everything after that is the machinery for finding out what a judge's number is worth. Five agreement numbers instead of one, because a constant judge beats the real one on exact agreement. The document in the prompt, which takes the judge's correlation with the person from 0.02 to 0.69 and puts the two systems the right way round. Fifty words of padding that raise 22 of 48 scores and lower none. A swap of two answers that flips half the pairwise verdicts. A calibration table that says the judge's 4 is a person's 3, chosen on half the set and held on the other. Four wordings of one rubric that put the correlation anywhere from 0.16 to 0.73. And a verdict that names the two uses the judge is fit for and refuses the third.
Two models run through Ollama on your own machine: qwen3:4b is the
judge and llama3.2:1b wrote the answers and serves as the small judge
lesson 2 measures. Every call is at seed 0 and temperature 0 through
Ollama's JSON output, so a score is an integer and a rerun gives the
same integer. The forty eight answers were recorded once and are written
into the setup cell, so the person's scores stay attached to the words
they were given for; only the judge runs live.
How this course teaches
Lesson 1 is a tour: six answers, two judges, the answer key graded, and a prediction about forty eight. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a number that is true and misleading. A judge that is better than chance and below a constant. A third of a person, measured on a number a constant scores half on. A deployable judge chosen because the good one used a reference nobody has, when the document does as well. A length bias fixed by a sentence that flattens the judge. A pairwise verdict that survived a swap and still ranks the systems backwards. A threshold that agrees 79 percent on the rows that chose it. A stable rerun taken as a settled number. A judge refused for every use because of the one it fails. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you will be able to do
- Build a judged set with a person's scores and a second reading, and measure the floor a constant judge sets.
- Split agreement into five numbers and say which one a constant can game and which it cannot.
- Put the gold answer and then the document into the judge's prompt and measure what the wrong answers and the system ranking do.
- Test a judge for length bias with the content held fixed, and price the rubric sentence that removes it.
- Judge every pair in both orders, count the flips, find the judge's lean, and hold the surviving verdicts to the person.
- Translate the judge's scale through a calibration table, and choose a pass threshold and a flag bar on one half of the set and hold them on the other.
- Measure what a rerun proves, what a hot seed moves, what a vote buys, and what one clause of the rubric is worth.
- Put ranking, flagging and grading on one table and write the verdict that names the uses the judge is fit for.
The lessons
1. A model grades a model. Six answers with a person's scores. The small judge gives five of them a 4. The big judge gives the right deposit a 4 and two wrong answers a 5, and marks its own answer key wrong nine times of twenty four. The predict cell runs it over all forty eight: correlation 0.02, and the worse system scores higher.
2. Forty eight answers and a rubric. Two systems, a person's rubric, a second reading made by a stated rule that differs on 6 of 48. The small judge scores 40 of 48 answers 4 and sits below a rule that always answers 1. The system gap the person sees, 1.3 points, the small judge sees as 0.04.
3. Agreement is five numbers. A judge that is one point harsh everywhere has near zero exact agreement and near perfect rank correlation. A constant has 42 percent exact agreement and no correlation at all. Kappa gives the constant exactly zero, the second reading 0.83, and the big judge 0.07.
4. The judge without the facts. Seventeen of twenty wrong answers scored 4 or 5 with no reference, none scored 5 with the gold answer, two scored 4 or 5 with the document. Correlation 0.02, then 0.69, then 0.69. The systems change places. The answer key scores 5 seventeen times once the judge is shown it.
5. Longer looks better. Fifty words of nothing take 27 words to 77 and move 22 of 48 scores up and none down; the wrong answers go from 3.9 to 4.8. The reference does not cure it. A sentence in the rubric does, by collapsing the free judge onto 4 and costing the referenced one 0.14 of correlation.
6. Which one came first. Swap two answers and 12 of 24 verdicts flip. The small judge leans towards the first slot, the big judge towards the second. Asking both orders leaves 11 verdicts, 7 of them the person's, and the pairwise count still ranks the systems backwards where the pointwise referenced score has them right.
7. What a four means. The judge's 4 is a person's 3.4 and its 3 a person's 1.5. A table built on twelve questions cuts the error on the other twelve by about a fifth. A pass mark of 4 chosen on one half holds at 83 percent on the other. A bar of 3 catches 23 of 26 bad answers with one false alarm. The free judge has no monotone table.
8. The same verdict twice. A rerun changes nothing. Hot seeds change 29, 27 and 8 rows, and two of them agree with each other on 46. A vote among them is no better than one cold run. Four rubric wordings under one document put the correlation between 0.16 and 0.73 and the system gap between 0.2 and 1.4; one clause, the anchor for a 1, is worth more than any sampler setting.
9. Three uses, one judge. On 500 draws of eight questions the referenced judge ranks the systems the person's way 439 times and the free judge 41. At a bar of 3 it catches 23 of 26 bad answers. On a single answer it matches the person 35 percent of the time against a second person's 88. Fit to rank, fit to flag, unfit to grade, and the document in every prompt.
Requirements
- Python 3.9 or later with
ollamaandpandas.
pip install ollama pandas
- A running Ollama (https://ollama.com) with two models pulled once:
ollama pull qwen3:4b
ollama pull llama3.2:1b
- EVAL 101 for the test set and the baseline, EVAL 201 for where human scores come from, PANDAS 101 for the tables.
Every lesson's setup cell checks for the server and both models and says the one command to run if anything is missing. The forty eight answers and the person's scores are written into the setup cell, so nothing is downloaded and the set is the same on every machine. Every number in the prose was produced by the cell above it, on this machine, at seed 0 and temperature 0; another Ollama build may score a row differently, and every check in the course is a relation or a band rather than a typed answer.
What to read
The Ollama documentation on structured outputs explains the format
argument every judge call uses, which is why a score is an integer and
no lesson parses a sentence. For the subject, search for LLM as a
judge with position bias and verbosity bias; the paper that
introduced the MT Bench comparisons measured both on much larger judges
than the one on your machine and found the shapes lessons 5 and 6 find
here. EVAL 201 in this series is where the human scores come from and
how much two people agree, which is the ceiling lesson 3 holds the judge
to. EVAL 301 is next: the thing being judged is an agent's whole run,
and a right answer that took forty tool calls is the failure.