TOOL 101
tool-101 · v1.0.0
ivyx✓
Tool Calling. Hand a model a function and watch the call come back as data, measure which of two overlapping tools one word in a description sends a question to, price a tool list in tokens, and decide what your handler says when the argument is valid and names nothing.
What this course is for
By the end of this course you can hand a model a function, read the call it sends back as data, run it and hand the result back, write the description that decides which of two overlapping tools is called, price a tool list in prompt tokens, refuse a call before it is dispatched, and decide what your handler returns when the argument is well formed and names nothing.
What you will be able to do
- Declare a tool by hand, read the name and the arguments back out of the reply, score both models on the tool they reach for, and find that picking the tool and calling it are two different scores
- Write the dispatcher, put a result back as a tool message, read the answer out of the second turn, and find the six questions one round trip cannot answer
- Put two tools that take the same argument in front of both models, move one word between their descriptions, and find which model it moves and which it does not
- Price three, six and thirteen tools in prompt tokens, count the wrong turns each list causes on both models, and find the cost certain and the damage conditional
- Count the argument faults on both models, find the small one returning the schema where the call belongs, and write the check that catches every shape of fault before dispatch
- Refuse a faulty call two ways, blankly and by naming the fault, count the recoveries on each model, and find that the sentence is worth everything above a floor and nothing below it
- Ask about cars that are not on the lot, watch a perfectly formed argument pass every check, and measure what an empty result, a null and a refusal each make the model do next
- Put the course into one rule, build a naive tool setup and a designed one, and measure both end to end on the same questions
Who it is for
Learners who have called a model and want to give it a function to call back, and who would rather measure what a tool description, a tool list and a handler's reply do than argue about them.
Before you start
- LLM 101, for the tokens every tool list is priced in
- PROMPT 101, for measuring what a change to a sentence does
Lesson path
A declaration, a call that has not run, and the round trip that answers
- 1A function the model can call40 min
Hand the model three functions, watch a call come back as data, run the function yourself, hand the result back, and predict what a small model sends instead
- 2The call is data and nothing ran55 min
Declare a tool by hand, read the name and the arguments back out of the reply, score both models on the tool they reach for, and find that picking the tool and calling it are two different scores
- 3Running it and handing the result back55 min
Write the dispatcher, put a result back as a tool message, read the answer out of the second turn, and find the six questions one round trip cannot answer
The sentence that routes a question, and what a longer tool list costs
- 4The description is the router55 min
Put two tools that take the same argument in front of both models, move one word between their descriptions, and find which model it moves and which it does not
- 5What the tool list costs55 min
Price three, six and thirteen tools in prompt tokens, count the wrong turns each list causes on both models, and find the cost certain and the damage conditional
The faults a schema cannot stop, what a refusal buys, and an argument that names nothing
- 6The schema asks, it does not fence55 min
Count the argument faults on both models, find the small one returning the schema where the call belongs, and write the check that catches every shape of fault before dispatch
- 7What the handler says back55 min
Refuse a faulty call two ways, blankly and by naming the fault, count the recoveries on each model, and find that the sentence is worth everything above a floor and nothing below it
- 8Valid, and naming nothing55 min
Ask about cars that are not on the lot, watch a perfectly formed argument pass every check, and measure what an empty result, a null and a refusal each make the model do next
The five clauses, a naive desk and a designed one, scored side by side
- 9A rule for handing over a tool60 min
Put the course into one rule, build a naive tool setup and a designed one, and measure both end to end on the same questions
About this course
TOOL 101 · Tool Calling
Hand a model three functions and ask how many kilometres a car has done, and what comes back is not an answer and not a lookup. It is a small piece of data naming one of your functions and filling in its arguments. Nothing has run. Your program decides whether to run it, runs it, hands the result back, and the answer comes out of the next turn.
That is the whole mechanism, and everything interesting is in the
three choices around it. How many tools you declare: three cost about
500 prompt tokens on every call and thirteen cost about 1,200, and the
longer list leaves the big model at 28 of 30 while the small one falls
from 23 to 10. What the descriptions say: with two tools taking the
same argument, moving the single word price between them moves three
of eight calls on the small model and none on the big one. And what
your handler returns: an empty object makes the model go looking
elsewhere and answer nobody on 8 of 8, while an empty string, which is
what dict.get(key, "") gives you, makes the strongest model in this
course invent a mileage for a car that does not exist on 8 of 8. A
sentence naming the absence answers all eight and invents nothing.
This course measures all of that against two local models through
Ollama, qwen3.5:9b with reasoning off as the caller that works and
llama3.2:1b as the caller that does not. Nothing needs an account, a
key or the network once the models are pulled.
Three modules and a judgment. The call: a declaration, a request that has not run, a dispatcher, and the six questions of thirty that one round trip answers on none of. The description: two tools the schema cannot separate, and what a tool list costs in tokens and in answers. The arguments: half the small model's calls carrying the declaration where the call belongs, what a careful refusal buys on each model, and an argument that passes every check and names nothing. Judgment: five clauses, a naive desk and a designed one, scored side by side.
How this course teaches
Lesson 1 is a tour: one question through the whole round trip, five pieces at a time, and one prediction. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: a proposal that sounds right, refused by computing the thing a second way rather than by argument.
- A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a measurement taken on the model least able to show the fault. A tool description validated on the strongest model available. A long tool list shipped because it cost nothing there, then read by a cheaper tier. A tool call scored on the name it chose and not on whether its arguments could be dispatched. A schema believed to be enforced somewhere. An error message written carefully for a caller that cannot act on it. And a handler returning the value its storage layer gave it. Every diagnose cell is one of those, and the refusal is always a second count rather than an opinion.
What you need
- A running Ollama with
qwen3.5:9bandllama3.2:1bpulled, about 8 GB together. The setup cell of every lesson checks and says what to do if one is missing; in IVYX Studio the LLMS panel installs Ollama and pulls the models. pip install ollama pandasin the kernel's environment.- A lesson makes between twenty and a hundred short calls. Most run in two to four minutes, and the big model answers a question in about two seconds.
Your machine may answer a question or two differently from the one the course was built on. The prose says which numbers are the built machine's, and every check holds a band or a relation that another Ollama build satisfies too.