LLM 101
llm-101 · v1.0.0
ivyx✓
Tokens, Context, Sampling. Count what a text costs by the model's own report, build the list the count comes from, measure which sampling knob changes an answer, and watch a window fill and a conversation forget.
What this course is for
By the end of this course you can price any text by the model's own token count and say why a word count is not a price, build the byte pair list a count comes from, choose the sampling settings that make a run repeatable, the model's own answer, or varied on purpose, tell a fact from an invention by counting answers, detect a silently cut prompt from the one number that reveals it, and budget a conversation so its first turn survives the window.
What you will be able to do
- Count tokens with the model's own report, and price the chat wrapper, digits, emoji and capitals against prose
- Build a byte pair encoder from a corpus, and watch frequency rather than meaning decide what a word costs
- Show that temperature zero is a rerun under any seed, and that a seeded hot run repeats without being the model's answer
- Measure that a change of seed moves more answers than doubling the temperature, and that top-p and top-k can make a hot run cold
- Separate facts from inventions by counting distinct answers under several seeds, and refuse a model as a random number generator
- Detect a cut prompt from the processed count against the sent count, and put what must survive at the tail
- Watch a conversation lose its first turn under a small window, and read what the model says when it has
- Budget a window as wrapper, instructions, history and reserved answer, and hold the number against a live conversation
Who it is for
Learners who have sent a prompt to a model and want to know what it costs, why the same prompt answers differently, and what happens to a long conversation, measured on a local model rather than explained.
Before you start
- PANDAS 101, for the tables every lesson prints
- Python basics, functions and dicts, at the level of PYTHON 101
Lesson path
What a text costs by the model's own count, and where the pieces come from
- 1The model does not read words40 min
Watch one paragraph cost three prices, one seed change an answer, and one conversation forget a name, then predict what Turkish costs
- 2Counting what a text costs55 min
Build the counter from the model's own report, price twenty texts with it, and find what the chat wrapper, digits and capitals cost
- 3Cutting text by hand55 min
Build a byte pair encoder from Northgate's documents, watch a word it saw become one piece while a misspelling stays five, and compare it to the model's list
Temperature, seed, top-p and top-k, and which one decides the answer
- 4Temperature zero is a rerun55 min
Ask ten facts at temperature zero twice and with two seeds and find nothing changes, then turn the heat up and find which answers become wrong
- 5The seed decides more than the temperature55 min
Hold the seed and move the temperature, hold the temperature and move the seed, and find which knob changes the answers; then find two settings that make a hot run cold
- 6Facts and inventions55 min
Count the distinct answers each prompt has under three hot seeds, separate the prompts with one answer from those with three, and find that a model asked for a random number is not a random number generator
What the model sees, what a full window drops, and what the model says about it
- 7The window fills silently55 min
Put a fact at the head of a prompt, add filler until the model answers from the filler, read the one number that reveals the drop, and move the fact to the tail
- 8The conversation forgets55 min
Tell the model your name, hold six ordinary exchanges under a small window, ask your name, and read what the model says when its first turn has been cut
The window as arithmetic, verified against a live conversation
- 9A budget for the window60 min
Write the window as a budget of wrapper, instructions, history and answer, compute how many exchanges fit, and hold the number against a live conversation
About this course
LLM 101 · Tokens, Context, Sampling
A paragraph about a car warranty costs 73 tokens in English, 90 in Turkish and 120 in German, for the same meaning. Two words sent as a chat message cost 27, because a wrapper you never typed goes around every message. Ask a small model the red planet at temperature zero and it says Mars every time on every seed; ask at 0.7 with the seed changed and it says Venus, and a run that repeats is not the same thing as the model's answer. Put a fact at the head of a prompt, add filler until the prompt no longer fits the window, and the model answers from the filler with no sign anything was dropped, except one number in the reply. Tell it your name, talk about tyres for six turns under a small window, and it says it has no information about your personal details.
This course measures all of that on llama3.2:1b through a local
Ollama. Nothing needs an account, a key or the network once the model
is pulled. Every number in the prose was produced by the cells, and
the checks hold shapes and bands where another machine may draw
differently.
Three modules and a judgment. Tokens: the counter from the model's own report, and a byte pair encoder built from seventy documents so that you have seen where a list comes from and why frequency, not meaning, sets a word's price. Sampling: temperature zero as a rerun, the seed moving more answers than doubling the temperature, top-p and top-k making a hot run cold, and facts sorted from inventions by counting answers, ending on the model asked for a random number and giving its favourite. The window: a cut detected from the processed count against the sent count, the tail surviving where the head does not, and a conversation losing its first turn. Judgment: the window as a budget of wrapper, instructions, history and reserved answer, computed before a conversation starts and held against a live one.
How this course teaches
Lesson 1 is a tour: the three surprises in six worked cells and one prediction. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a plausible number. A word count read as a price. An encoder learned from your own documents read as the model's list. A seeded run that repeats read as the model's stable answer. Heat read as variety. A wrong answer read as the model's memory limit when the prompt was cut before it arrived. The model's sentence that it has no information read as a fact about the conversation. A window divided by the exchange cost with the wrapper and the answer left out. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you need
- A running Ollama with
llama3.2:1bpulled (about 1.3 GB). The setup cell of every lesson checks and says what to do if it is missing; in IVYX Studio the LLMS panel installs Ollama and pulls the model. pip install ollama pandasin the kernel's environment.- A machine that runs a 1B model at a few tokens a second is enough. A lesson makes between twenty and a hundred short calls and runs in under a minute.
Your machine may draw differently from the one the course was built on. The prose says which numbers are the built machine's, and every check holds a relation or a band that another Ollama build satisfies too.