RAG 301
rag-301 · v1.0.0
ivyx✓
Structured and Multi-Source Retrieval: query a table, parse a codebase, walk a graph that joins them, route a question to the source that can answer it, and report the disagreement instead of merging it.
What this course is for
By the end of this course you can answer a question from a table, a codebase and a document set together, say which source each part of the answer came from, and report a disagreement between two sources instead of quietly picking one.
What you will be able to do
- Treat a query as a retriever whose result is rows, and catch the value the question's word does not match
- Carry the row count an aggregate was computed over, and tell an empty result from a refusal
- Parse a codebase with ast into a symbol index, and tell a definition from the lines that mention it
- Build a graph that joins a table to a document set, and walk it in both directions
- Route a question to a source by its shape, and measure that against routing by vocabulary
- Give every answer a source, a locator and a verdict, and keep the attempt that refused
- Normalise two type systems before comparing them, and separate a conflict from an absence
- Date two editions of a policy, read the edition the code declares, and report a superseded implementation
Who it is for
Learners who finished RAG 101 and SQL 101 and want to answer questions from a table, a codebase and a document set at once, with a record that says where every part of the answer came from.
Before you start
- RAG 101, for the keyword retriever this course reuses as its document source
- SQL 101, for DuckDB and the queries the table source is made of
- PYTHON 102, for modules and docstrings, which the code source is parsed for
Lesson path
A table, a document set and a codebase, each asked the same question
- 1Three answers40 min
Ask one question of a table, a document set and a codebase, and get three different answers
- 2The table as a retriever55 min
Treat a query as a retriever whose result is rows, and meet the value the question's word does not match
- 3Symbols, not text55 min
Parse the code instead of searching it, and tell a definition from the ten lines that mention it
The edges between the sources, the choice of which to ask, and the record of what answered
- 4A graph you can walk55 min
Build the edges that join the table to the documents and answer what neither holds alone
- 5Routing by question shape55 min
Pick the source from the question's shape, measure it, and watch vocabulary routing lose
- 6The provenance record55 min
Give every answer a source, a locator and a verdict, including the source that could not answer
Two sources about one fact, and two editions of one policy
- 7The contradiction55 min
Compare two sources field by field, normalise the types first, and separate a conflict from an absence
- 8The stale edition55 min
Date every answer by the edition it implements, and report code that applies a superseded policy
Three service desk calls that need every source at once
- 9Answer it from all three60 min
Answer a question that needs the table, the code and the documents, and print the contradiction instead of a merged number
About this course
RAG 301 · Structured and Multi-Source Retrieval
Every course in this branch so far has had one source: 52 documents,
chunked, embedded, reranked and compiled. A company does not keep its
answers in one place. Northgate Motors keeps the lot in a table, because
a table is what you count and sort. It keeps the paperwork as documents,
because a policy is prose. And it runs a sales system, which is code, and
code is the only one of the three that is executed rather than read. Ask
all three how long the warranty is and you get 24 months from the current
policy, 12 from the edition it replaced, and 12 from warranty.py, whose
docstring says it implements the 2023 edition.
This course builds a retriever over each of them, a graph that joins the
table to the documents, a router that picks between them by the shape of
the question, and a record that says where every part of an answer came
from. Then it spends two lessons on what to do when two sources
disagree, which is to report it rather than merge it. No model is
involved: the sources are DuckDB, Python's own ast and RAG 101's
keyword retriever, and the queries are written by hand, because turning a
question into SQL is a model's job and another course's subject.
How this course teaches
Lesson 1 is a tour: it asks one question of three sources, follows a two hop path to an answer none of them holds alone, and shows a comparison reporting 25 disagreements where there is one. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
- Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
- A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.
No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.
The particular danger of this subject is an answer that is true about the
system and false about the question. gearbox = 'automatic' returns no
rows, which reads as a fact about the lot. An average over a column with
a blank in it is about six cars when the question was about seven. A sort
answers the lowest recorded mileage and leaves two cars out of the
ranking entirely. A text search over code returns a docstring in the
wrong module as the definition. A two hop walk lands on the right
document and reads the wrong sentence of it. A comparison between a
boolean column and the words yes and no disagrees 24 times out of 24.
A fallback onto a second source always answers, including on a question it
has nothing about. Every diagnose cell is one of those, and every cross
check is the second route that refuses it.
What you will be able to do
- Treat a query as a retriever, check a column's own vocabulary before filtering on it, and carry the row count an aggregate was computed over.
- Tell an empty result from a refusal, and give both a shape a caller can act on.
- Parse a codebase into a symbol index with
ast, read a constant without importing the module that holds it, and list the functions that read it. - Build a graph that joins a table to a document set, walk it in both directions, and encode a policy rule as edges.
- Route a question to a source by its shape, score a router against labelled questions, and measure the vocabulary alternative rather than arguing about it.
- Give every answer a source, a locator and a verdict, and keep the attempt that refused beside the one that answered.
- Normalise two type systems before comparing them, and separate a conflict from an absence.
- Date two editions of one policy, read the edition a module declares for itself, and report a superseded implementation with the line numbers and the functions that read it.
The lessons
1. Three answers. One question, three sources: a BinderException
from the table, two warranty editions from the documents with the
superseded one ranked first, and 12 from warranty.py line 6. The oil
interval for A-103, two hops away and in no single document. A-114's row
against its own description. And 25 disagreements in 72 comparisons,
24 of them caused by the type DuckDB chose.
2. The table as a retriever. gearbox = 'automatic' returns 0 rows
and 'auto' returns 11. The average price of a Torres is 17,166.67 over
six of seven rows, or 14,714.29 counting the blank. ORDER BY km answers
A-116 at 12,000 and never ranks the two cars with no reading. A missing
column raises and an unmatched filter does not, and a retriever that
folds those together reports both as nothing found.
3. Symbols, not text. The word warranty is on 10 lines in 3 of the
6 files; the length is defined once, at warranty.py line 6. 35 symbols,
16 functions and 19 constants, each with a file and a line. A name search
that keeps short words returns five hits because is_collector and
trade_in_value match on is and in. ast.literal_eval reads a
constant without running its module.
4. A graph you can walk. 63 nodes, 60 edges, four relations from the
table and one built by running the collector rule over the year column.
A-103 to Nube to svc-nube: every 20,000 km. The same two hops for the
1971 Adler land on the right document and read the wrong sentence, which
is what a correct path cannot protect you from.
5. Routing by question shape. Five patterns, 17 of 17 on the single
source questions. Routing by each source's own vocabulary: 5 of 13, and
nearly everything lands on the documents. A schema guard that fixes the
two questions the table cannot answer breaks the two whose words are
mileage and odometer while the column is called km. The router
returns a list after that, because a route is a guess until a source
answers.
6. The provenance record. One shape over four sources: a query, a file and a line, a path, a document id. Two questions are refused by the table and answered by the documents on the second attempt, and both answers are wrong, one from the superseded warranty edition and one from a FAQ about reserving a car. The record does not fix that; it makes it visible in a column.
7. The contradiction. 72 comparisons, 25 disagreements, 2 after the boolean column is normalised. One conflict, A-114, whose table row says damage history and whose description says no accident history, with the company's own FAQ as the evidence that those are the same claim. One absence, A-123, priced only in its description. The merge fills the absence and refuses the conflict.
8. The stale edition. Two editions, one of which says it replaces the other. All four warranty constants in the code match the superseded edition exactly and none matches the current one, and the module's docstring says which edition it implements. Three of the four are read by a function that decides a customer's claim. The retriever ranks the superseded edition first on every warranty question, and on one of them it does not return the current edition at all.
9. Answer it from all three. Three service desk calls. The code
refuses a claim the 2025 edition covers, because it applies the 2023
limits. The code would pay a claim on the 1971 collector car, because
in_warranty never asks what kind of car it is and that clause exists
only in the policy. The third claim, a 2024 sale, both refuse, which is
the control. Two defects, two line numbers, and a sentence you write.
Requirements
- Python 3.9 or later with
duckdbandpandas.
pip install duckdb pandas
- RAG 101, for the keyword retriever this course reuses as its document source; SQL 101, for DuckDB and the queries; PYTHON 102 for modules and docstrings, which lesson 3 parses.
No model, no network and no API key. Every lesson's setup cell rewrites
all three sources from scratch: the table as w1_cars.csv, the 52
documents as w1_docs.csv, and six Python files under w1_repo, so
editing them is safe and rerunning the cell restores them. Every number
in the prose was produced by the cell above it, and every check is a
relation or a count rather than a typed answer.
What to read
DuckDB's documentation on reading CSV files explains type inference,
which is the mechanism behind lesson 1's surprise and lesson 7's first
run: a column of yes and no read as BOOLEAN. Python's ast module
documentation is the reference for lesson 3; the node types and
ast.get_docstring are the two pages to read. For the subject this
course deliberately does not do, the term of art is text to SQL, and
Spider and BIRD are the benchmarks for turning a question into a query
with a model. RAG 101's lesson 6 is where the stale warranty edition
first outranked the current one, and this course is what finally settles
it, with a date rather than a score.