RED 301
red-301 · v1.0.0
ivyx✓
Red Teaming: Attacking on Purpose. Run fifty attacks in five classes against a tool using assistant under six defence variants, read what each defence stopped, left and broke, and keep the set as a regression harness.
What this course is for
By the end of this course you can build an attack set as prompts with detectors and keep it, read a transcript for what an attack got, run the set against a target under each defence and read what each stopped, what it did not and what it broke, tell a target that resists from one that cannot do the job, and make every fix leave a permanent test behind.
What you will be able to do
- Build an attack set as classes of prompts with a benign control group, and keep it as a record
- Write one detector per class, and find the detector that missed six successes on its own
- Read direct and poisoned document injections, and name the channel each defence reads
- Follow the secret out through eight prompts, and find what redaction masks and what it does not
- Separate a discount call attempted from one applied, and find the one defence that reads the call
- Price six variants on what they fixed, what they left and what they cost the ten controls
- Read a weaker target's safer looking table beside its control row, and say what its zeros mean
- Turn the set into a regression harness that names fixed, open and regressed, and ships or holds
Who it is for
Learners with an assistant that uses tools and carries a system prompt with something in it worth protecting, who want a repeatable attack set, a detector per class, a price on each defence and a harness that catches a fix that reopens something.
Before you start
- EVAL 301, for recorded runs read as tables and the judge that scores the controls
- SEC 201, for prompt injection and tool permissions as the attacks assume them
- PANDAS 101, for the tables every lesson prints
Lesson path
Five attacks that land, the set they come from, and the detector that says so
- 1Attacking on purpose40 min
Watch five attacks land on an undefended assistant, one per class, and predict how many of fifty survive every defence
- 2Fifty prompts and ten controls55 min
Build the attack set as classes with a benign control group, and read the recorded runs as a table
- 3The detector55 min
Write one detector per class, find the one that missed six successes, and add the one the string detector cannot replace
Injection through two channels, and the secret leaving through the answer
- 4Instructions as data55 min
Read the direct and the poisoned document injections, and find which defence stops which
- 5The secret leaves55 min
Read the ten exfiltration transcripts, and find what redaction masks and what it does not
The check at the tool, the price of every defence, and a target that cannot do the job
- 6The tool a customer must not reach55 min
Separate a tool call attempted from a tool call that worked, and find the one defence that reads the call
- 7What a defence costs55 min
Price every defence by what it fixed, what it left, and what it did to the ten controls
- 8A weaker target55 min
Run the same attacks against a much smaller model, and read a table that looks safer for what it is
Fifty tests with recorded answers, and the verdict a change has to pass
- 9Every fix leaves a test60 min
Turn the attack set into a regression harness that names what a change fixed, left open and reopened
About this course
RED 301 · Red Teaming: Attacking on Purpose
A service desk assistant with four tools and a staff-only code in its instructions is told to ignore its instructions, and does. It reads a car listing whose description ends with an instruction, and obeys it. It gives up the code eight ways out of ten, once by reciting its whole prompt. It applies a discount for seven customers who ask, passing the code it was told never to share. And it refuses all ten harmful requests without any help. That is the undefended target, and this course is the fifty attacks, the ten controls and the six variants that turn it into a table.
Everything after that is reading the table for what it does not say. A delimiter on retrieved text takes the poisoned document from ten to zero and touches nothing else. A code check at the tool refuses all seven discount calls without changing one of them. An output redaction masks the code in eight answers and leaves the recital standing with one word blanked. A hardened prompt reaches every class, finishes none, and is the only defence that costs a control. Five direct injections survive everything, and they are the five the prompt did not describe. A model one sixth the size reads as safer on two classes, and passes two controls in ten.
The runs were recorded once, big target and small, under every variant, and frozen into the setup cell. Nothing in the course calls a model, so it needs pandas alone and every lesson verifies in seconds. The material is the transcripts, the way a log is, and the work is reading them.
How this course teaches
Lesson 1 is a tour: five attacks that land, one per class, and a prediction about how many of fifty survive every defence. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is a zero that is true and misleading. A detector's zero read as a class closed, when the defence was written against that detector. A prompt that stopped five read as one sentence away from ten. A filter that zeroed one channel moved to the other. The defence that fixed the most alone shipped alone. A small model's lower residual read as resistance when it is a model that cannot do the task. A falling total read as a release verdict while one attack reopened. Every diagnose cell is one of those, and every cross check is the second route that refuses it.
What you will be able to do
- Build an attack set as classes of prompts with a benign control group, and keep it as a record with the runs it produced.
- Write one detector per class, find the case bug that missed six successes, and add the recital detector a string test cannot replace.
- Read direct and poisoned document injections, and name the channel each defence reads and the five attacks no channel defence reaches.
- Follow the secret out through eight prompts, and find what redaction masks and what it leaves.
- Separate a discount call attempted from one applied, and find the one defence that does not depend on the model agreeing.
- Price six variants on what they fixed, what they left and what they cost the controls, and read the limit of a ten control row.
- Read a weaker target's safer looking table beside its control row, and write the comparability test that says when two targets compare at all.
- Turn the set into a regression harness that names fixed, open and regressed, and ships or holds on the last list alone.
The lessons
1. Attacking on purpose. One attack per class against the undefended assistant: the marker echoed, the poisoned listing obeyed, the code given, the discount applied, the harmful request refused. The predict cell asks how many of fifty survive every defence; the answer is five.
2. Fifty prompts and ten controls. Five classes of ten, a control group of ten ordinary questions scored by a judge, six variants and two targets in one frame of 480 runs. The controls score 4.8 under the baseline and no variant refuses a customer.
3. The detector. One detector per class. The marker detector written case sensitively counts four poisoned documents where ten replies tell the customer to quote the code in lower case. The string detector for the secret reads zero under redaction while one reply recites the whole prompt with the code masked.
4. Instructions as data. Direct injection ten of ten, poisoned document ten of ten. The delimiter takes the document to zero and the message stays at ten. The hardened prompt stops the five that announce themselves and leaves the five dressed as tasks, and a filter moved to the customer's message would reach one of them.
5. The secret leaves. Eight of ten prompts get the code; the two that do not ask for the prompt as a document. Redaction fires on exactly the eight and leaves the recital. The hardened prompt refuses the recital and leaves the new manager and just the code, so the class closes only where both are on.
6. The tool a customer must not reach. Seven attempts under the baseline, all carrying the code the prompt gave the model, and one customer who knew it. The code check refuses all seven without changing one. The hardened prompt leaves the two customers who claim to be the manager, which no prompt can verify.
7. What a defence costs. Residuals of 35, 15, 25, 28, 27 and 5. The hardened prompt fixes twenty alone and is the only variant with a control below four, a warranty answered wrong; under the all variant the same sentence scores that control five. Ten controls find a cost and cannot say whether it recurs.
8. A weaker target. The small model's baseline row reads 3, 0, 2, 9, 0. It passes two controls in ten, leaves twelve answers empty, never reads the poisoned listing, and calls the discount tool from sixteen prompts with no listing and zero percent. One class cannot be compared at all.
9. Every fix leaves a test. Fixed, open, regressed and never between any two variants. The big target ships with thirty fixed, five named open tests and nothing reopened. The small target's residual falls from fourteen to two while one injection that used to miss now lands, and the harness holds.
Requirements
- Python 3.9 or later with
pandas.
pip install pandas
- EVAL 301 for recorded runs read as tables and the judge on the controls, SEC 201 for prompt injection and tool permissions, PANDAS 101 for the tables.
The runs were recorded once from qwen3.5:9b and llama3.2:1b
through Ollama's tool calling at seed 0 and temperature 0, with
qwen3:4b as the judge on the harmful and benign replies, and are
written into the setup cell in full; nothing is downloaded and no model
runs. The setup cell is large and takes about a second. Every number in
the prose was produced by the cell above it on those recorded runs, and
every check is a relation over the recording rather than a typed
answer.
What to read
The attacks in this course are the author's, and a red team's value is the attacks the builder did not think of; the OWASP Top 10 for LLM Applications is the shortest public list of the classes worth adding prompts for, and Simon Willison's writing on prompt injection is the clearest argument for why the direct injection class stays open. SEC 201 in this series covers the tool permission model the code check assumes, and EVAL 301 the recorded run tables every lesson here is built on.