← All courses
IVYXSTUDIO · COURSE

EVAL 301

eval-301 · v1.0.0

ivyx✓

Agent Evals: read a hundred and fifty recorded agent runs as routes, find the right answers no tool produced, count the wrong and late calls, price every step, and pick an agent under a budget and a task mix.

advanced485 min9 lessonsen
#evaluation#agents#trajectory#tool-calls#cost#ai-series#advanced

What this course is for

By the end of this course you can read an agent run as a record of steps, tell a right answer the tools produced from one that came from nowhere, count the wrong, malformed and redundant calls, price a run from its tokens, say how a run ended and what a budget does to it, compose a trajectory score that separates agents the answer score cannot, and pick an agent for a task mix and a budget in writing.

What you will be able to do

  • Read a recorded run as steps, calls and results, and build a table of counts a log gives for free
  • Test every tool result for the task's fact and split each right answer into grounded and lucky
  • Count calls to the wrong tool, calls with empty arguments, and answers written as calls that reached no tool
  • Count the calls made after the fact was in hand and the exact repeats, and the calls a run needed
  • Price every step of a run from its tokens, watch the prompt grow, and find the bill that is not in the calls
  • Tell a finish from a plain reply, measure what an instruction about ending cost, and cap the quantity that grows
  • Compose one pass or fail from six criteria, and show which criterion a ranking hangs on
  • Choose an agent under a budget, a task mix and a latency limit, and write the decision with its arguments beside it

Who it is for

Learners who have an agent loop calling tools and a score on its final answers, and want to know what the route cost, whether the answers came from the tools, and which agent to ship under a budget.

Before you start

  • EVAL 202, for the judge whose scores sit on every final answer here
  • AGENT 201, for the loop these runs were recorded from
  • COST 201, for the price table every run is priced with

Lesson path

The run3 lessons

A run as a record, a table of what it did, and where each right answer came from

  1. 1Twenty nine of thirty, twice40 min

    Watch two agents reach the same score by different routes, and predict which of five costs most per right answer

  2. 2A run is a record55 min

    Read a recorded run as steps, calls and results, build the run table, and score five agents on their answers alone

  3. 3Right answer, wrong route55 min

    Find the right answers no tool result contained, and hold the key, the judge and the tools against each other

What the trajectory scores3 lessons

The calls that could not have worked, the calls that were not needed, and what every step cost

  1. 4The wrong tool and the malformed call55 min

    Count calls to the wrong tool, calls with no arguments and calls that never reached a tool, and find the agent that made forty

  2. 5A call after the fact was in hand55 min

    Count the calls made after the answer was already on screen, and the exact repeats, and score a run on them

  3. 6Cost per run55 min

    Price every step of every run, watch the context grow, and find the agent whose cost is in the tokens it thinks with

Scoring2 lessons

How a run ends, what a budget does, and one verdict composed from the whole route

  1. 7How a run ends55 min

    Tell a finish from a plain reply, measure what a prompt demanding finish cost, and put a budget on a run instead of a step cap

  2. 8The trajectory score55 min

    Compose one pass or fail from the answer, the route and the cost, and watch the ranking change

Judgment1 lesson

Five agents, a budget, a task mix and a decision that carries its arguments

  1. 9Pick the agent60 min

    Put five agents on one table with a budget and a task mix, and write the verdict that names the one to ship and why

About this course

EVAL 301 · Agent Evals: Trajectory, Tool Calls, Cost

Two agents answer 29 of 30 service desk questions correctly. The judge from EVAL 202 cannot tell them apart. One of them makes a third more tool calls, reads the full document after the search had already shown the answer twice as often, and costs a quarter more per right answer. The agent that makes the fewest calls of all five is the most expensive per run, because its bill is in the twelve hundred tokens of reasoning it writes between calls. And twelve of the hundred and fifty right answers in the set were never in any tool result: the agent knew, or guessed, and was right this time.

This course scores the route, not the answer. It holds one hundred and fifty recorded agent runs, five agents on the same thirty tasks about Northgate's documents and listings with the same five tools, every step's calls, arguments, results, tokens and seconds frozen into the setup cell, with EVAL 202's referenced judge's score on each final answer. Nothing in the course calls a model, so it needs pandas alone and every lesson verifies in seconds. The material is the runs, the way a log is, and the work is reading them.

How this course teaches

Lesson 1 is a tour: four runs read step by step, two agents tied on the answer and apart on the route, a right answer about the wrong car, and sixteen calls in one step. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.

  • A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
  • Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
  • An exercise that is broken when you open it.
  • A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
  • A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.

No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.

The particular danger of this subject is a number that is true and misleading. Two agents tied on the answer and called interchangeable. A free score that agrees with the judge nine times in ten, on the agents where it never matters. A call count read as diligence when most of the calls could not have worked. Reading in full read as care when the same model without it scores the same. Fewest calls read as cheapest. A slow agent given a higher step cap it never came near. A pass count read as quality when the ranking reverses without one criterion. A best row shipped outright with no budget beside it. Every diagnose cell is one of those, and every cross check is the second route that refuses it.

What you will be able to do

  • Read a recorded run as steps, calls and results, and build a table of the counts a log gives for free.
  • Test every tool result for the task's fact and split each right answer into grounded and lucky.
  • Count calls to the wrong tool, calls with empty arguments, and answers written as calls that reached no tool.
  • Count the calls made after the fact was in hand and the exact repeats, and the calls a run actually needed.
  • Price every step from its tokens, watch the prompt grow across a run, and find the bill that is not in the calls.
  • Tell a finish from a plain reply, measure what an instruction about ending cost, and cap the quantity that grows.
  • Compose one pass or fail from six criteria and show which criterion a ranking hangs on.
  • Choose an agent under a budget, a task mix and a latency limit, and write the decision with its arguments beside it.

The lessons

1. Twenty nine of thirty, twice. The big agent's four step run on a listing question, the thorough agent's same task with one more read, the small agent answering about the wrong car, and sixteen calls in one step. The predict cell prices the five agents per right answer.

2. A run is a record. Steps, calls, results, tokens and endings. The run table of one hundred and fifty rows. On answers alone big and thorough tie at 29, the thinker 23, the small agent 19, the strict small agent 6; on the record the tied pair are 16 calls apart.

3. Right answer, wrong route. The key matches words, the judge reads meaning, the run knows the source. Twelve right answers no tool produced, five the small agent's and four the thinker's. The key is safe on the two big agents and on nobody else.

4. The wrong tool and the malformed call. The three larger models make no malformed call in ninety runs. The small agent makes six. Told it must always call finish, it makes forty two, writes fifteen answers as JSON, and finishes exactly as often as before.

5. A call after the fact was in hand. Thorough makes 35 calls after the fact was on screen to big's 17 for the same answers, and the extra reading bought no answer the record can find. Exact repeats live in one run: thirteen of the strict agent's burst.

6. Cost per run. A four step run reads 3,659 tokens where four first steps would read 2,712. The thinker costs twice big per run and nearly three times per right answer with half the calls, and four fifths of its bill is what it wrote.

7. How a run ends. No run reached the eight step cap; the longest was four steps. A 0.02 budget cuts nineteen of the thinker's runs and brings the small agent level with thorough on right answers kept.

8. The trajectory score. Six criteria, one verdict: big 23, thorough 16, small 14, thinker 9, strict 2. Drop the budget and the thinker climbs two places; drop the repeat criterion and nobody moves. Every agent fails the route on the two hop tasks.

9. Pick the agent. A fitness table, a floor per agent, a task mix, and a chooser that takes a budget and a latency limit. Big above about 0.015 a run, the small agent below it with a floor of zero on listing questions, thorough never, the thinker never.

Requirements

  • Python 3.9 or later with pandas.
pip install pandas
  • EVAL 202 for the judge whose scores sit on every answer, AGENT 201 for the loop the runs came from, COST 201 for the price table.

The runs were recorded once from llama3.2:1b, qwen3.5:9b and qwen3:4b through Ollama's tool calling at seed 0 and temperature 0, and are written into the setup cell in full; nothing is downloaded and no model runs. The setup cell is large and takes about a second. Every number in the prose was produced by the cell above it on those recorded runs, and every check is a relation over the recording rather than a typed answer.

What to read

The Ollama documentation on tool calling explains the shape every run was recorded from: a tools list on the request, tool_calls on the reply, and a tool role message carrying each result back. For the subject, search for agent evaluation with trajectory and tool use; the benchmark papers that score agents on browsing and software tasks all separate the final answer from the route, and the metrics here are the plain versions of theirs. EVAL 202 in this series is where the judge on every answer was measured, COST 201 is where the price table comes from, and OBS 301 is next: the same measurements on traffic while it happens rather than on a recording.