← All courses
IVYXSTUDIO · COURSE

RAG 301

rag-301 · v1.0.0

ivyx✓

Structured and Multi-Source Retrieval: query a table, parse a codebase, walk a graph that joins them, route a question to the source that can answer it, and report the disagreement instead of merging it.

advanced485 min9 lessonsen
#retrieval#rag#sql#duckdb#ast#graph#provenance#ai-series#advanced

What this course is for

By the end of this course you can answer a question from a table, a codebase and a document set together, say which source each part of the answer came from, and report a disagreement between two sources instead of quietly picking one.

What you will be able to do

  • Treat a query as a retriever whose result is rows, and catch the value the question's word does not match
  • Carry the row count an aggregate was computed over, and tell an empty result from a refusal
  • Parse a codebase with ast into a symbol index, and tell a definition from the lines that mention it
  • Build a graph that joins a table to a document set, and walk it in both directions
  • Route a question to a source by its shape, and measure that against routing by vocabulary
  • Give every answer a source, a locator and a verdict, and keep the attempt that refused
  • Normalise two type systems before comparing them, and separate a conflict from an absence
  • Date two editions of a policy, read the edition the code declares, and report a superseded implementation

Who it is for

Learners who finished RAG 101 and SQL 101 and want to answer questions from a table, a codebase and a document set at once, with a record that says where every part of the answer came from.

Before you start

  • RAG 101, for the keyword retriever this course reuses as its document source
  • SQL 101, for DuckDB and the queries the table source is made of
  • PYTHON 102, for modules and docstrings, which the code source is parsed for

Lesson path

Three sources3 lessons

A table, a document set and a codebase, each asked the same question

  1. 1Three answers40 min

    Ask one question of a table, a document set and a codebase, and get three different answers

  2. 2The table as a retriever55 min

    Treat a query as a retriever whose result is rows, and meet the value the question's word does not match

  3. 3Symbols, not text55 min

    Parse the code instead of searching it, and tell a definition from the ten lines that mention it

Joining them3 lessons

The edges between the sources, the choice of which to ask, and the record of what answered

  1. 4A graph you can walk55 min

    Build the edges that join the table to the documents and answer what neither holds alone

  2. 5Routing by question shape55 min

    Pick the source from the question's shape, measure it, and watch vocabulary routing lose

  3. 6The provenance record55 min

    Give every answer a source, a locator and a verdict, including the source that could not answer

Disagreement2 lessons

Two sources about one fact, and two editions of one policy

  1. 7The contradiction55 min

    Compare two sources field by field, normalise the types first, and separate a conflict from an absence

  2. 8The stale edition55 min

    Date every answer by the edition it implements, and report code that applies a superseded policy

Judgment1 lesson

Three service desk calls that need every source at once

  1. 9Answer it from all three60 min

    Answer a question that needs the table, the code and the documents, and print the contradiction instead of a merged number

About this course

RAG 301 · Structured and Multi-Source Retrieval

Every course in this branch so far has had one source: 52 documents, chunked, embedded, reranked and compiled. A company does not keep its answers in one place. Northgate Motors keeps the lot in a table, because a table is what you count and sort. It keeps the paperwork as documents, because a policy is prose. And it runs a sales system, which is code, and code is the only one of the three that is executed rather than read. Ask all three how long the warranty is and you get 24 months from the current policy, 12 from the edition it replaced, and 12 from warranty.py, whose docstring says it implements the 2023 edition.

This course builds a retriever over each of them, a graph that joins the table to the documents, a router that picks between them by the shape of the question, and a record that says where every part of an answer came from. Then it spends two lessons on what to do when two sources disagree, which is to report it rather than merge it. No model is involved: the sources are DuckDB, Python's own ast and RAG 101's keyword retriever, and the queries are written by hand, because turning a question into SQL is a model's job and another course's subject.

How this course teaches

Lesson 1 is a tour: it asks one question of three sources, follows a two hop path to an answer none of them holds alone, and shows a comparison reporting 25 disagreements where there is one. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.

  • A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
  • Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
  • An exercise that is broken when you open it.
  • A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
  • A challenge that ends in a record or a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.

No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.

The particular danger of this subject is an answer that is true about the system and false about the question. gearbox = 'automatic' returns no rows, which reads as a fact about the lot. An average over a column with a blank in it is about six cars when the question was about seven. A sort answers the lowest recorded mileage and leaves two cars out of the ranking entirely. A text search over code returns a docstring in the wrong module as the definition. A two hop walk lands on the right document and reads the wrong sentence of it. A comparison between a boolean column and the words yes and no disagrees 24 times out of 24. A fallback onto a second source always answers, including on a question it has nothing about. Every diagnose cell is one of those, and every cross check is the second route that refuses it.

What you will be able to do

  • Treat a query as a retriever, check a column's own vocabulary before filtering on it, and carry the row count an aggregate was computed over.
  • Tell an empty result from a refusal, and give both a shape a caller can act on.
  • Parse a codebase into a symbol index with ast, read a constant without importing the module that holds it, and list the functions that read it.
  • Build a graph that joins a table to a document set, walk it in both directions, and encode a policy rule as edges.
  • Route a question to a source by its shape, score a router against labelled questions, and measure the vocabulary alternative rather than arguing about it.
  • Give every answer a source, a locator and a verdict, and keep the attempt that refused beside the one that answered.
  • Normalise two type systems before comparing them, and separate a conflict from an absence.
  • Date two editions of one policy, read the edition a module declares for itself, and report a superseded implementation with the line numbers and the functions that read it.

The lessons

1. Three answers. One question, three sources: a BinderException from the table, two warranty editions from the documents with the superseded one ranked first, and 12 from warranty.py line 6. The oil interval for A-103, two hops away and in no single document. A-114's row against its own description. And 25 disagreements in 72 comparisons, 24 of them caused by the type DuckDB chose.

2. The table as a retriever. gearbox = 'automatic' returns 0 rows and 'auto' returns 11. The average price of a Torres is 17,166.67 over six of seven rows, or 14,714.29 counting the blank. ORDER BY km answers A-116 at 12,000 and never ranks the two cars with no reading. A missing column raises and an unmatched filter does not, and a retriever that folds those together reports both as nothing found.

3. Symbols, not text. The word warranty is on 10 lines in 3 of the 6 files; the length is defined once, at warranty.py line 6. 35 symbols, 16 functions and 19 constants, each with a file and a line. A name search that keeps short words returns five hits because is_collector and trade_in_value match on is and in. ast.literal_eval reads a constant without running its module.

4. A graph you can walk. 63 nodes, 60 edges, four relations from the table and one built by running the collector rule over the year column. A-103 to Nube to svc-nube: every 20,000 km. The same two hops for the 1971 Adler land on the right document and read the wrong sentence, which is what a correct path cannot protect you from.

5. Routing by question shape. Five patterns, 17 of 17 on the single source questions. Routing by each source's own vocabulary: 5 of 13, and nearly everything lands on the documents. A schema guard that fixes the two questions the table cannot answer breaks the two whose words are mileage and odometer while the column is called km. The router returns a list after that, because a route is a guess until a source answers.

6. The provenance record. One shape over four sources: a query, a file and a line, a path, a document id. Two questions are refused by the table and answered by the documents on the second attempt, and both answers are wrong, one from the superseded warranty edition and one from a FAQ about reserving a car. The record does not fix that; it makes it visible in a column.

7. The contradiction. 72 comparisons, 25 disagreements, 2 after the boolean column is normalised. One conflict, A-114, whose table row says damage history and whose description says no accident history, with the company's own FAQ as the evidence that those are the same claim. One absence, A-123, priced only in its description. The merge fills the absence and refuses the conflict.

8. The stale edition. Two editions, one of which says it replaces the other. All four warranty constants in the code match the superseded edition exactly and none matches the current one, and the module's docstring says which edition it implements. Three of the four are read by a function that decides a customer's claim. The retriever ranks the superseded edition first on every warranty question, and on one of them it does not return the current edition at all.

9. Answer it from all three. Three service desk calls. The code refuses a claim the 2025 edition covers, because it applies the 2023 limits. The code would pay a claim on the 1971 collector car, because in_warranty never asks what kind of car it is and that clause exists only in the policy. The third claim, a 2024 sale, both refuse, which is the control. Two defects, two line numbers, and a sentence you write.

Requirements

  • Python 3.9 or later with duckdb and pandas.
pip install duckdb pandas
  • RAG 101, for the keyword retriever this course reuses as its document source; SQL 101, for DuckDB and the queries; PYTHON 102 for modules and docstrings, which lesson 3 parses.

No model, no network and no API key. Every lesson's setup cell rewrites all three sources from scratch: the table as w1_cars.csv, the 52 documents as w1_docs.csv, and six Python files under w1_repo, so editing them is safe and rerunning the cell restores them. Every number in the prose was produced by the cell above it, and every check is a relation or a count rather than a typed answer.

What to read

DuckDB's documentation on reading CSV files explains type inference, which is the mechanism behind lesson 1's surprise and lesson 7's first run: a column of yes and no read as BOOLEAN. Python's ast module documentation is the reference for lesson 3; the node types and ast.get_docstring are the two pages to read. For the subject this course deliberately does not do, the term of art is text to SQL, and Spider and BIRD are the benchmarks for turning a question into a query with a model. RAG 101's lesson 6 is where the stale warranty edition first outranked the current one, and this course is what finally settles it, with a date rather than a score.