← All courses
IVYXSTUDIO · COURSE

EVAL 202

eval-202 · v1.0.0

ivyx✓

LLM as Judge: hold a judge model to answers a person already scored, give it the facts, measure its length and position bias, translate its scale, and say in writing which uses it is fit for.

intermediate485 min9 lessonsen
#evaluation#llm-as-judge#rubric#calibration#bias#ollama#ai-series#intermediate

What this course is for

By the end of this course you can write a judge rubric and hold the judge to answers a person already scored, read five agreement numbers and say which one hides an offset, find the facts a judge needs in its prompt, measure length and position bias and undo what can be undone, translate the judge's scale into a person's, and say which uses a judge is fit for before you trust a number from it.

What you will be able to do

  • Build a judged set with a person's scores and a second reading, and measure the floor a constant judge sets
  • Split agreement into exact, within one, error, rank correlation and kappa, and say what each one cannot see
  • Put the gold answer and then the document into the judge's prompt, and measure what the wrong answers and the system ranking do
  • Pad every answer with content free text, count the scores that rose, and price the rubric sentence that stops it
  • Judge every pair in both orders, count the flips, find each judge's lean, and hold the surviving verdicts to the person
  • Translate the judge's scale through a calibration table, and choose a pass threshold and a flag bar on one half and hold them on the other
  • Measure a rerun, three hot seeds and a majority vote, and the range a rubric's wording moves the number across
  • Put ranking, flagging and grading on one table and write the verdict that names the uses the judge is fit for

Who it is for

Learners who have a model producing free text answers and want a second model to score them, and who need to know before trusting the score what it correlates with, what moves it, and which decisions it can carry.

Before you start

  • EVAL 101, for the test set, the scorer, the baseline and the regression check this course builds on
  • EVAL 201, for where human scores come from and how much two people agree
  • PANDAS 101, for the tables every lesson prints

Lesson path

The judge3 lessons

Two judges, a set a person scored, and five ways to say how much they agree

  1. 1A model grades a model40 min

    Watch two judges score six answers a person already scored, and predict how the better one does over forty eight

  2. 2Forty eight answers and a rubric55 min

    Build the judged set with its human scores, write the rubric, and find the judge that answers one number

  3. 3Agreement is five numbers55 min

    Measure exact, within one, error, rank correlation and kappa, and find the one that hides an offset

What the judge reads3 lessons

The facts it was never shown, the padding it rewards, and the slot it prefers

  1. 4The judge without the facts55 min

    Put the document in the prompt and watch the wrong answers fall and the systems change places

  2. 5Longer looks better55 min

    Pad every answer with nothing and count the scores that rose, then measure what the fix costs

  3. 6Which one came first55 min

    Judge every pair in both orders, count the flips, and find the lean each judge has

Calibration2 lessons

What the judge's four means, and what moves it

  1. 7What a four means55 min

    Map the judge's scale onto the person's, pick a threshold on half the set and hold it on the other half

  2. 8The same verdict twice55 min

    Measure a rerun, a hot seed and a majority vote, then change one clause of the rubric and watch the ranking move

Judgment1 lesson

Three uses, two judges, one table and a verdict

  1. 9Three uses, one judge60 min

    Put ranking, flagging and grading on one table and write the verdict that says which the judge is fit for

About this course

EVAL 202 · LLM as Judge

A service desk answer is a sentence, and is this a good answer has no label to match it against. The usual move is to hand it to a second model with a rubric and take the score. This course holds that score to forty eight answers a person already scored, and the first thing it finds is that the judge gives two confidently wrong answers a 5, ranks the worse of two systems above the better, and marks its own answer key wrong nine times out of twenty four.

Everything after that is the machinery for finding out what a judge's number is worth. Five agreement numbers instead of one, because a constant judge beats the real one on exact agreement. The document in the prompt, which takes the judge's correlation with the person from 0.02 to 0.69 and puts the two systems the right way round. Fifty words of padding that raise 22 of 48 scores and lower none. A swap of two answers that flips half the pairwise verdicts. A calibration table that says the judge's 4 is a person's 3, chosen on half the set and held on the other. Four wordings of one rubric that put the correlation anywhere from 0.16 to 0.73. And a verdict that names the two uses the judge is fit for and refuses the third.

Two models run through Ollama on your own machine: qwen3:4b is the judge and llama3.2:1b wrote the answers and serves as the small judge lesson 2 measures. Every call is at seed 0 and temperature 0 through Ollama's JSON output, so a score is an integer and a rerun gives the same integer. The forty eight answers were recorded once and are written into the setup cell, so the person's scores stay attached to the words they were given for; only the judge runs live.

How this course teaches

Lesson 1 is a tour: six answers, two judges, the answer key graded, and a prediction about forty eight. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.

  • A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
  • Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
  • An exercise that is broken when you open it.
  • A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
  • A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.

No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.

The particular danger of this subject is a number that is true and misleading. A judge that is better than chance and below a constant. A third of a person, measured on a number a constant scores half on. A deployable judge chosen because the good one used a reference nobody has, when the document does as well. A length bias fixed by a sentence that flattens the judge. A pairwise verdict that survived a swap and still ranks the systems backwards. A threshold that agrees 79 percent on the rows that chose it. A stable rerun taken as a settled number. A judge refused for every use because of the one it fails. Every diagnose cell is one of those, and every cross check is the second route that refuses it.

What you will be able to do

  • Build a judged set with a person's scores and a second reading, and measure the floor a constant judge sets.
  • Split agreement into five numbers and say which one a constant can game and which it cannot.
  • Put the gold answer and then the document into the judge's prompt and measure what the wrong answers and the system ranking do.
  • Test a judge for length bias with the content held fixed, and price the rubric sentence that removes it.
  • Judge every pair in both orders, count the flips, find the judge's lean, and hold the surviving verdicts to the person.
  • Translate the judge's scale through a calibration table, and choose a pass threshold and a flag bar on one half of the set and hold them on the other.
  • Measure what a rerun proves, what a hot seed moves, what a vote buys, and what one clause of the rubric is worth.
  • Put ranking, flagging and grading on one table and write the verdict that names the uses the judge is fit for.

The lessons

1. A model grades a model. Six answers with a person's scores. The small judge gives five of them a 4. The big judge gives the right deposit a 4 and two wrong answers a 5, and marks its own answer key wrong nine times of twenty four. The predict cell runs it over all forty eight: correlation 0.02, and the worse system scores higher.

2. Forty eight answers and a rubric. Two systems, a person's rubric, a second reading made by a stated rule that differs on 6 of 48. The small judge scores 40 of 48 answers 4 and sits below a rule that always answers 1. The system gap the person sees, 1.3 points, the small judge sees as 0.04.

3. Agreement is five numbers. A judge that is one point harsh everywhere has near zero exact agreement and near perfect rank correlation. A constant has 42 percent exact agreement and no correlation at all. Kappa gives the constant exactly zero, the second reading 0.83, and the big judge 0.07.

4. The judge without the facts. Seventeen of twenty wrong answers scored 4 or 5 with no reference, none scored 5 with the gold answer, two scored 4 or 5 with the document. Correlation 0.02, then 0.69, then 0.69. The systems change places. The answer key scores 5 seventeen times once the judge is shown it.

5. Longer looks better. Fifty words of nothing take 27 words to 77 and move 22 of 48 scores up and none down; the wrong answers go from 3.9 to 4.8. The reference does not cure it. A sentence in the rubric does, by collapsing the free judge onto 4 and costing the referenced one 0.14 of correlation.

6. Which one came first. Swap two answers and 12 of 24 verdicts flip. The small judge leans towards the first slot, the big judge towards the second. Asking both orders leaves 11 verdicts, 7 of them the person's, and the pairwise count still ranks the systems backwards where the pointwise referenced score has them right.

7. What a four means. The judge's 4 is a person's 3.4 and its 3 a person's 1.5. A table built on twelve questions cuts the error on the other twelve by about a fifth. A pass mark of 4 chosen on one half holds at 83 percent on the other. A bar of 3 catches 23 of 26 bad answers with one false alarm. The free judge has no monotone table.

8. The same verdict twice. A rerun changes nothing. Hot seeds change 29, 27 and 8 rows, and two of them agree with each other on 46. A vote among them is no better than one cold run. Four rubric wordings under one document put the correlation between 0.16 and 0.73 and the system gap between 0.2 and 1.4; one clause, the anchor for a 1, is worth more than any sampler setting.

9. Three uses, one judge. On 500 draws of eight questions the referenced judge ranks the systems the person's way 439 times and the free judge 41. At a bar of 3 it catches 23 of 26 bad answers. On a single answer it matches the person 35 percent of the time against a second person's 88. Fit to rank, fit to flag, unfit to grade, and the document in every prompt.

Requirements

  • Python 3.9 or later with ollama and pandas.
pip install ollama pandas
ollama pull qwen3:4b
ollama pull llama3.2:1b
  • EVAL 101 for the test set and the baseline, EVAL 201 for where human scores come from, PANDAS 101 for the tables.

Every lesson's setup cell checks for the server and both models and says the one command to run if anything is missing. The forty eight answers and the person's scores are written into the setup cell, so nothing is downloaded and the set is the same on every machine. Every number in the prose was produced by the cell above it, on this machine, at seed 0 and temperature 0; another Ollama build may score a row differently, and every check in the course is a relation or a band rather than a typed answer.

What to read

The Ollama documentation on structured outputs explains the format argument every judge call uses, which is why a score is an integer and no lesson parses a sentence. For the subject, search for LLM as a judge with position bias and verbosity bias; the paper that introduced the MT Bench comparisons measured both on much larger judges than the one on your machine and found the shapes lessons 5 and 6 find here. EVAL 201 in this series is where the human scores come from and how much two people agree, which is the ceiling lesson 3 holds the judge to. EVAL 301 is next: the thing being judged is an agent's whole run, and a right answer that took forty tool calls is the failure.