DATA 202
data-202 · v1.0.0
ivyx✓
Vector and Hybrid Store Operations: watch two models answer one corpus differently, measure exact against approximate recall, find the filter that eats the results, meet the loud model failure and the silent one, reindex without going dark, and put size, latency and recall on one page.
What this course is for
By the end of this course you can choose an index type against a measured recall number rather than a reputation, filter on metadata without silently losing results, reindex after the embedding model changes without going dark, and say what a hybrid retriever costs against what it buys.
What you will be able to do
- Query one corpus under two embedding models and watch the top result change with no error
- Build exact search, measure recall at k against the gold set, and name the ceiling every later lesson is compared to
- Trade recall for speed with an approximate index and measure both halves of the trade
- Compare post filter and pre filter and count the answers each returns empty for a narrow filter
- Partition an index by a field and measure the search it buys and the coverage it forecloses
- Tell the loud model failure at a different width from the silent one at the same width, caught only by recall
- Reindex by building beside and swapping atomically rather than in place, which is up and wrong
- Measure a hybrid against its parts and put size, latency, recall and rebuild cost on one page
Who it is for
Learners who finished RAG 201 and RAG 202 and now have to operate the store they built: choose its index, filter it, reindex it, and defend the choice.
Before you start
- RAG 201, for chunking, the two local embedding models, and the dense retriever this course operates
- RAG 202, for the hybrid fusion this course measures again and the gold set it scores against
Lesson path
Two stores of one corpus, exact search, and the approximate trade
- 1The neighbour that moved40 min
Query one chunk set under two embedding models and watch the top result change with no error anywhere
- 2Exact search, and what it costs55 min
Build brute force search yourself, measure it against the gold set, and establish the recall ceiling every later lesson is compared to
- 3Approximate search55 min
Trade recall for speed on purpose, measure recall@k against exact, and name the parameter that moved it
A filter that ate the results, and a partition that forecloses queries
- 4The filter that ate the results55 min
Compare post filter and pre filter over the gold set at three values of k and count the empty answers
- 5Metadata as part of the key55 min
Partition by a field, measure what partitioning buys and what it forecloses
The model changing underneath the index, and reindexing without going dark
- 6The model changed underneath you55 min
Meet the loud failure at a different width, then the silent one at the same width, and catch the second with the gold set
- 7Reindex without going dark55 min
Build the new index beside the old one, swap, and describe exactly what a reader sees during the swap
The hybrid measured against its parts, and the store on one page
- 8Hybrid, measured55 min
Fuse BM25 and vectors over one corpus and compare against rag-202's published numbers on the same questions
- 9Operate it60 min
Put size, latency, recall and rebuild cost on one page and make a choice you can defend with the page
About this course
DATA 202 · Vector and Hybrid Store Operations
RAG 201 built a dense retriever and RAG 202 measured what its context did to the answer. This course is about the vector store itself, the thing between the corpus and the retriever, and the operations that keep it honest: choosing an index type against a measured recall number rather than a reputation, filtering without silently losing results, partitioning by a field, surviving a model change, reindexing without going dark, and fusing a keyword ranking into a vector one. Each is a place a store quietly gives a wrong answer while looking like it is working, and this course measures every one.
Every number is measured live in the kernel against Northgate's 52 documents,
chunked as RAG 201 chunks them and embedded twice, once with nomic-embed-text at
768 dimensions and once with all-minilm at 384, both local and both already
pulled by the wave one courses. The 24 gold questions, each labelled with the
document that answers it, are what make recall a number rather than an impression,
and the lot table's brands, attached to the listing chunks, are what a filter has
to filter on. Nothing here leaves the machine once the two models are pulled.
How this course teaches
Lesson 1 is a tour: it builds two stores of one corpus, asks all 24 questions of each, and lets you predict how many top results move between the two models. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.
- A prediction you commit to before the cell runs, graded on the reasoning and not the guess, where being wrong is the point.
- Warmups: a one line blank or a short exercise under the theory it practices, each with a four rung hint ladder whose last rung explains and still does not hand over the code.
- An exercise that is broken when you open it.
- A diagnose cell that runs, prints a confident and plausible conclusion, and is wrong, with the evidence that refuses it already on the screen.
- A challenge that ends in a sentence you write, which the tutor grades.
No cell in this course passes in the state it ships, and that is checked mechanically before the course is published.
The particular danger of this subject is a store that answers confidently and wrong. A post-filter returns an empty list where the caller asked for five. A same-width model swap returns five neighbours and they are the wrong five, with no error. An in-place reindex answers every request from a half-built index. A naive hybrid scores one question below the vector search it replaced. Every diagnose cell is one of those, and the only instrument that catches any of them is recall measured against the gold set.
What you will be able to do
- Query one corpus under two embedding models and watch the top result change with no error anywhere.
- Build exact search, measure recall at k against the gold set, and name the ceiling every later lesson is compared to.
- Trade recall for speed with an approximate index and measure both halves of the trade.
- Compare post filter and pre filter and count the answers each returns empty for a narrow filter.
- Partition an index by a field and measure the search it buys and the coverage it forecloses.
- Tell the loud model failure at a different width from the silent one at the same width, caught only by recall.
- Reindex by building beside and swapping atomically rather than in place, which is up and wrong.
- Measure a hybrid against its parts and put size, latency, recall and rebuild cost on one page.
The lessons
1. The neighbour that moved. Two stores of one corpus, nomic and all-minilm. The top chunk differs for 4 of 24 questions with no error, and the two models put the gold document first for 21 and 18, a quality difference only the gold set sees.
2. Exact search, and what it costs. One dot product against every chunk. Recall is 21, 23, 23, 24 for nomic at k of 1, 5, 10, 20, 18, 22, 23, 24 for all-minilm, rising fast and flattening. 24 is the ceiling, and the cost is 190 chunks times 768 dimensions per query.
3. Approximate search. Random-hyperplane buckets. At 4 bits recall@5 is 11 of 24 against exact's 23, examining about 11 candidates instead of 190; at 8 bits 5 and 3; at 12 bits 2 and under 1. Recall and candidates fall together as the one knob, the number of planes, rises.
4. The filter that ate the results. A brand post-filter returns empty for 22 of 24 questions at k of 5, because the narrow brand rarely appears in an unfiltered top five; pre-filter returns zero empties. A bigger k chips at the empties and never removes them.
5. Metadata as part of the key. Partitioning by brand buys a search over a fifth of the store and a structural isolation, and forecloses coverage: only 2 of 24 questions have a branded answer, so a brand partition cannot answer the other 22.
6. The model changed underneath you. A different width, 384 against 768, raises a ValueError, the loud safe failure. A same-width swap, stood in for by a rotation, errors nowhere and collapses recall@5 from 23 to 1, caught only by the gold set.
7. Reindex without going dark. In place, the index is a mixture of two spaces and recall dips for the whole rebuild while answering every request, up and wrong. Beside, readers stay on the old index at recall 23 and the swap is atomic, no dip.
8. Hybrid, measured. Dense 21, BM25 16, reciprocal rank fusion 20, one below dense. The fusion loses q03, the part-exchange question dense owns and BM25 never ranks, and no fusion constant recovers a document one ranker is blind to.
9. Operate it. Size, query cost, recall and rebuild cost on one page. No row wins every column, so a choice needs a constraint: a recall floor of 22 picks all-minilm as the cheapest that clears it, defended by the whole page.
Requirements
Python 3.10 or later with ollama, numpy and pandas:
pip install ollama numpy pandas
A running Ollama with two embedding models pulled, which the setup cell checks:
ollama pull nomic-embed-text
ollama pull all-minilm
Both are a few hundred megabytes and never need the network after that; in IVYX Studio the LLMS panel installs and pulls them. Embedding the corpus twice takes a few seconds the first time and is cached after, so the whole course runs in about twenty seconds of model time.
What to read outside this course
The model cards for nomic-embed-text and all-minilm on the Ollama library give
the dimensions and token windows the course measures. Any vector database's guide
to index types, exact against approximate, and to metadata filtering explains the
machinery this course builds in numpy so you can see it; the course's contribution
is to make you measure recall before trusting a number. RAG 201 lesson 8 raced
these two models and RAG 202 measured the hybrid this course measures again, so
both are the direct prerequisites.