PROMPT 101
prompt-101 · v1.0.0
ivyx✓
Writing Instructions. Measure a role, a constraint, three examples and a format instruction on one extraction task with an answer key, find the parts do not add, price each prompt per right answer, and choose on half the set and confirm on the other.
What this course is for
By the end of this course you can score a prompt against an answer key with two readers and read its misses before changing it, measure what a role, examples, a format instruction and a constraint each do on the model that will read them, show that the parts do not add by running every combination, price a prompt per right answer, and choose one on half the set and confirm it on the other.
What you will be able to do
- Build a task with an answer key, a strict parser and a run that keeps every answer and its token counts
- Add a facts reader and counts of refusals and echoes, so a score says how a prompt failed
- Measure four role wordings and find each turns most of the set into refusals on a small model and none on a bigger one
- Count the answers that copy the examples instead of answering, and find the layout that stops it
- Measure the format instruction, its refusals, what the constraint adds beside it, and what the order of the lines does
- Run all sixteen combinations of the parts and show the additive prediction wrong by twenty or more
- Price every prompt in tokens per right answer and find the best score is not the best prompt
- Choose a prompt on half the set, confirm it on the other half, and break the tie by price
Who it is for
Learners who have written prompts by following advice and want to know what each piece of advice does on the model they use, measured on an extraction task with an answer key rather than judged by eye.
Before you start
- LLM 101, for the token counts every price in lesson 8 uses
- PANDAS 101, for the tables every lesson prints
Lesson path
A parser, a facts reader and counts of refusals and echoes, so a score can say how a prompt failed
- 1Five prompts, one task40 min
Watch the same extraction asked five ways on three messages, see a role line turn into a refusal, and predict which part of a prompt moves the score over thirty
- 2A task with an answer key55 min
Build the prompt, the parser and the scorer, score the bare instruction at 24 of 30, and find that its six misses all read a model name as the model year
- 3Two scoreboards55 min
Build a reader that finds the facts in any shape of answer, score the strict constraint at 0 by the parser and 29 by the facts, and keep both scores from then on
The role, the examples, the format instruction and the order, each measured alone on this model
- 4The role that refuses55 min
Measure four role wordings and find every one turns most of the set into refusals, move the role into the system message, and ask a bigger model that never refuses
- 5Examples the model continues55 min
Count the answers that repeat an example instead of answering, find one example worse than three, find a cue that stops the copying, and find examples after the message trigger refusals
- 6The format instruction55 min
Measure the one part that raised the score alone, its four refusals, what the constraint adds beside it, and what the order of the parts does
All sixteen combinations, the failure of the additive prediction, and a price per right answer
- 7Parts are not additive55 min
Run all sixteen combinations of the four parts, predict each from the single parts' effects, and find the prediction wrong by twenty or more on most of them
- 8What a prompt costs55 min
Price every prompt in tokens per correct answer, and find the best scoring prompts cost twice what a prompt one answer behind them costs
A choice made on one half of the set and confirmed on the other
- 9Choose on half, confirm on the other60 min
Rank the sixteen prompts on the first fifteen messages, confirm the candidates on the last fifteen, break the tie by price, and write the verdict a shipped prompt needs
About this course
PROMPT 101 · Writing Instructions
One sentence of instruction and a customer's message score 24 of 30
on an extraction task with an answer key. Put a role line in front,
you are a service desk clerk at Northgate Motors, and the small
model refuses all thirty. Add a strict format constraint and it
scores 0 by the parser and 29 by the facts, because it answered in
its own shape. Add three worked examples and it scores 15, copying
the examples back on the other fifteen. Add a JSON format instruction
and it scores 26. Put all four together and it scores 30. This course
is those measurements, on llama3.2:1b through a local Ollama, and
the method that produced them.
Every claim about a prompt in the course is a difference between two counts, and both counts come from the cells. The task is thirty service desk messages naming a car's make, model year and mileage; the answer key is the three facts a person wrote down; the score is how many of thirty a prompt gets right, read by a parser. No judge, no opinion.
Three modules and a judgment. The task and the score: the parser, a second reader that finds the facts in any shape, and counts of refusals and echoes, because a score alone cannot say how a prompt failed. The parts: the role that turns an extraction into a refusal and does not on a bigger model, the examples the model continues instead of following and the layout that stops it, the format instruction that fixed the shape and made four customers look suspicious, and the order of the lines. Together: all sixteen combinations of the four parts, where the additive prediction is wrong by twenty or more on nine of fifteen, and a price per right answer that puts the best prompts at twice the cost of one a single answer behind. Judgment: choose on half the set, confirm on the other, break the tie by price.
How this course teaches
Lesson 1 is a tour: five prompts on three messages and one prediction. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a guide. Roles help; examples help; be strict; be clear. Each of those is measured here on one model and one task, and most of them are false there and true elsewhere, which is the lesson: a part of a prompt has no effect of its own, only an effect in a prompt, on a model, on a task, and the count is the only way to know it. Every diagnose cell is a guide's rule applied without measuring, and every cross check is the measurement.
What you need
- A running Ollama with
llama3.2:1bpulled (about 1.3 GB), andqwen3:4b(about 2.5 GB) for one comparison in lesson 4. The setup cell of every lesson checks and says what to do if a model is missing; in IVYX Studio the LLMS panel installs Ollama and pulls the models. pip install ollama pandasin the kernel's environment.- A lesson makes between thirty and five hundred short calls. Most run in a minute or two; the grid lessons take about four minutes each.
Your machine may answer a message or two differently from the one the course was built on. The prose says which numbers are the built machine's, and every check holds a band or a relation that another Ollama build satisfies too.