EVAL 301
eval-301 · v1.0.0
ivyx✓
Agent Evals: read a hundred and fifty recorded agent runs as routes, find the right answers no tool produced, count the wrong and late calls, price every step, and pick an agent under a budget and a task mix.
What this course is for
By the end of this course you can read an agent run as a record of steps, tell a right answer the tools produced from one that came from nowhere, count the wrong, malformed and redundant calls, price a run from its tokens, say how a run ended and what a budget does to it, compose a trajectory score that separates agents the answer score cannot, and pick an agent for a task mix and a budget in writing.
What you will be able to do
- Read a recorded run as steps, calls and results, and build a table of counts a log gives for free
- Test every tool result for the task's fact and split each right answer into grounded and lucky
- Count calls to the wrong tool, calls with empty arguments, and answers written as calls that reached no tool
- Count the calls made after the fact was in hand and the exact repeats, and the calls a run needed
- Price every step of a run from its tokens, watch the prompt grow, and find the bill that is not in the calls
- Tell a finish from a plain reply, measure what an instruction about ending cost, and cap the quantity that grows
- Compose one pass or fail from six criteria, and show which criterion a ranking hangs on
- Choose an agent under a budget, a task mix and a latency limit, and write the decision with its arguments beside it
Who it is for
Learners who have an agent loop calling tools and a score on its final answers, and want to know what the route cost, whether the answers came from the tools, and which agent to ship under a budget.
Before you start
- EVAL 202, for the judge whose scores sit on every final answer here
- AGENT 201, for the loop these runs were recorded from
- COST 201, for the price table every run is priced with
Lesson path
A run as a record, a table of what it did, and where each right answer came from
- 1Twenty nine of thirty, twice40 min
Watch two agents reach the same score by different routes, and predict which of five costs most per right answer
- 2A run is a record55 min
Read a recorded run as steps, calls and results, build the run table, and score five agents on their answers alone
- 3Right answer, wrong route55 min
Find the right answers no tool result contained, and hold the key, the judge and the tools against each other
The calls that could not have worked, the calls that were not needed, and what every step cost
- 4The wrong tool and the malformed call55 min
Count calls to the wrong tool, calls with no arguments and calls that never reached a tool, and find the agent that made forty
- 5A call after the fact was in hand55 min
Count the calls made after the answer was already on screen, and the exact repeats, and score a run on them
- 6Cost per run55 min
Price every step of every run, watch the context grow, and find the agent whose cost is in the tokens it thinks with
How a run ends, what a budget does, and one verdict composed from the whole route
- 7How a run ends55 min
Tell a finish from a plain reply, measure what a prompt demanding finish cost, and put a budget on a run instead of a step cap
- 8The trajectory score55 min
Compose one pass or fail from the answer, the route and the cost, and watch the ranking change
Five agents, a budget, a task mix and a decision that carries its arguments
- 9Pick the agent60 min
Put five agents on one table with a budget and a task mix, and write the verdict that names the one to ship and why
About this course
EVAL 301 · Agent Evals: Trajectory, Tool Calls, Cost
Two agents answer 29 of 30 service desk questions correctly. The judge from EVAL 202 cannot tell them apart. One of them makes a third more tool calls, reads the full document after the search had already shown the answer twice as often, and costs a quarter more per right answer. The agent that makes the fewest calls of all five is the most expensive per run, because its bill is in the twelve hundred tokens of reasoning it writes between calls. And twelve of the hundred and fifty right answers in the set were never in any tool result: the agent knew, or guessed, and was right this time.
This course scores the route, not the answer. It holds one hundred and fifty recorded agent runs, five agents on the same thirty tasks about Northgate's documents and listings with the same five tools, every step's calls, arguments, results, tokens and seconds frozen into the setup cell, with EVAL 202's referenced judge's score on each final answer. Nothing in the course calls a model, so it needs pandas alone and every lesson verifies in seconds. The material is the runs, the way a log is, and the work is reading them.
How this course teaches
Lesson 1 is a tour: four runs read step by step, two agents tied on the answer and apart on the route, a right answer about the wrong car, and sixteen calls in one step. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a number that is true and misleading. Two agents tied on the answer and called interchangeable. A free score that agrees with the judge nine times in ten, on the agents where it never matters. A call count read as diligence when most of the calls could not have worked. Reading in full read as care when the same model without it scores the same. Fewest calls read as cheapest. A slow agent given a higher step cap it never came near. A pass count read as quality when the ranking reverses without one criterion. A best row shipped outright with no budget beside it. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you will be able to do
- Read a recorded run as steps, calls and results, and build a table of the counts a log gives for free.
- Test every tool result for the task's fact and split each right answer into grounded and lucky.
- Count calls to the wrong tool, calls with empty arguments, and answers written as calls that reached no tool.
- Count the calls made after the fact was in hand and the exact repeats, and the calls a run actually needed.
- Price every step from its tokens, watch the prompt grow across a run, and find the bill that is not in the calls.
- Tell a finish from a plain reply, measure what an instruction about ending cost, and cap the quantity that grows.
- Compose one pass or fail from six criteria and show which criterion a ranking hangs on.
- Choose an agent under a budget, a task mix and a latency limit, and write the decision with its arguments beside it.
The lessons
1. Twenty nine of thirty, twice. The big agent's four step run on a listing question, the thorough agent's same task with one more read, the small agent answering about the wrong car, and sixteen calls in one step. The predict cell prices the five agents per right answer.
2. A run is a record. Steps, calls, results, tokens and endings. The run table of one hundred and fifty rows. On answers alone big and thorough tie at 29, the thinker 23, the small agent 19, the strict small agent 6; on the record the tied pair are 16 calls apart.
3. Right answer, wrong route. The key matches words, the judge reads meaning, the run knows the source. Twelve right answers no tool produced, five the small agent's and four the thinker's. The key is safe on the two big agents and on nobody else.
4. The wrong tool and the malformed call. The three larger models make no malformed call in ninety runs. The small agent makes six. Told it must always call finish, it makes forty two, writes fifteen answers as JSON, and finishes exactly as often as before.
5. A call after the fact was in hand. Thorough makes 35 calls after the fact was on screen to big's 17 for the same answers, and the extra reading bought no answer the record can find. Exact repeats live in one run: thirteen of the strict agent's burst.
6. Cost per run. A four step run reads 3,659 tokens where four first steps would read 2,712. The thinker costs twice big per run and nearly three times per right answer with half the calls, and four fifths of its bill is what it wrote.
7. How a run ends. No run reached the eight step cap; the longest was four steps. A 0.02 budget cuts nineteen of the thinker's runs and brings the small agent level with thorough on right answers kept.
8. The trajectory score. Six criteria, one verdict: big 23, thorough 16, small 14, thinker 9, strict 2. Drop the budget and the thinker climbs two places; drop the repeat criterion and nobody moves. Every agent fails the route on the two hop tasks.
9. Pick the agent. A fitness table, a floor per agent, a task mix, and a chooser that takes a budget and a latency limit. Big above about 0.015 a run, the small agent below it with a floor of zero on listing questions, thorough never, the thinker never.
Requirements
- Python 3.9 or later with
pandas.
pip install pandas
- EVAL 202 for the judge whose scores sit on every answer, AGENT 201 for the loop the runs came from, COST 201 for the price table.
The runs were recorded once from llama3.2:1b, qwen3.5:9b and
qwen3:4b through Ollama's tool calling at seed 0 and temperature 0,
and are written into the setup cell in full; nothing is downloaded and
no model runs. The setup cell is large and takes about a second. Every
number in the prose was produced by the cell above it on those
recorded runs, and every check is a relation over the recording rather
than a typed answer.
What to read
The Ollama documentation on tool calling explains the shape every run
was recorded from: a tools list on the request, tool_calls on the
reply, and a tool role message carrying each result back. For the
subject, search for agent evaluation with trajectory and tool
use; the benchmark papers that score agents on browsing and software
tasks all separate the final answer from the route, and the metrics
here are the plain versions of theirs. EVAL 202 in this series is where
the judge on every answer was measured, COST 201 is where the price
table comes from, and OBS 301 is next: the same measurements on traffic
while it happens rather than on a recording.