← All courses
IVYXSTUDIO · COURSE

OBS 301

obs-301 · v1.0.0

ivyx✓

Live Quality Monitoring: sample a stream, run an online judge, set an alarm with a measured false alarm rate, find the shift the mean hides, and tell a system change from a world change by what else moved.

advanced485 min9 lessonsen
#evaluation#monitoring#observability#drift#alarms#ollama#ai-series#advanced

What this course is for

By the end of this course you can turn a stream of answers into a daily quality series with a judge over a sample, set a baseline and an alarm rule and measure the rule's own false alarm rate, find the shift the mean hides by slicing, size the sample from the row variance, tell a system change from a world change by what else moved, choose the judge's reference so it can see the world move, and write the monitor down as a design.

What you will be able to do

  • Read a production stream as rows with a day, a topic, an answer and its telemetry, and measure what a normal week looks like
  • Draw a seeded daily sample, judge it live, and measure how far a sample's mean sits from the census
  • Set a baseline from one week, write a threshold rule and a cumulative one, and measure each rule's false alarm rate on a clean week
  • Find a shift the daily mean turns into a rumour and make it a fact by slicing to the questions it touched
  • Size the sample from the row variance and measure detection, delay and misses as rates over redraws
  • Read the telemetry columns across a shift and tell a system change from a world change
  • Judge against the system's own source and against the truth, and show which one cannot see the world move
  • Write a monitor as a design with its reasons, a cost, a measured false alarm rate and a runbook line per signature

Who it is for

Learners with an AI feature answering real traffic and a test set score they no longer trust, who want to know when the quality moved, by how much, in which slice, and whether it was the system or the world.

Before you start

  • EVAL 202, for the judge that scores every row and what its reference must be
  • EVAL 201, for where a labelled set comes from and how much people agree
  • PANDAS 101, for the tables and the groupbys every lesson uses

Lesson path

The stream3 lessons

Three weeks of traffic, a normal week measured, and a judge over a sample of each day

  1. 1Good on the test set40 min

    Watch a system that scored well on the test set drift through three weeks of traffic, and predict which days a monitor would catch

  2. 2Three weeks of traffic55 min

    Read the stream as rows with a day, a topic, an answer and a bill, and measure what a normal week looks like

  3. 3The online judge55 min

    Draw ten rows a day, judge them live, and measure how far a sample's daily mean sits from the census

Detecting3 lessons

A baseline, two rules, the shift the mean hides, and the rows a shift needs

  1. 4The baseline and the alarm55 min

    Set a baseline from week one, write a threshold rule and a cumulative one, and measure the rule's own false alarm rate

  2. 5The shift that hides55 min

    Find the week two shift the daily mean turns into a rumour, and make it a fact by slicing

  3. 6Sample size55 min

    Sweep rows per day and measure detection, delay and misses against what the row variance predicts

Attribution2 lessons

What else moved, and what the judge was comparing against

  1. 7What else moved55 min

    Read the telemetry columns across both shifts and tell a system change from a world change

  2. 8The judge's reference55 min

    Judge the shifted questions against the system's own source and against the truth, and watch one of them miss

Judgment1 lesson

The monitor as a design with its reasons, and the runbook

  1. 9Be the first to know60 min

    Write the monitor as a design, sample, rule, slices and reference, and the runbook line for each alarm

About this course

OBS 301 · Live Quality Monitoring in Production

A system that scored 3.7 on its test set answers three weeks of real traffic. On day 8 the dealership stops delivering cars and stops offering finance, nobody updates the documents, and one customer in eight is told about a delivery fee and a deposit that no longer exist. On day 15 retrieval silently stops. Nobody announces either. This course is about seeing both from a sample of ten rows a day, and about the difference between them: the first moves the daily mean by a quarter of a point and no telemetry column, the second by half a point and every column.

The material is 630 recorded rows, thirty a day for twenty one days, each a customer question answered by llama3.2:1b over the passages a keyword search returned, with EVAL 202's referenced judge score on every row against the truth of that day and against the document the system retrieved. Lesson 3 draws ten rows a day and runs the judge live; every other lesson reads the recorded census and draws samples from it, so a monitor's noise, false alarm rate and detection rate can be measured over a hundred redraws in seconds.

How this course teaches

Lesson 1 is a tour: a normal week, the whole census a monitor never has, the two events, and a ten row monitor asked what it would have seen. The eight lessons after it are graded work, each built the same way, and nine of their cells are yours.

  • A prediction you commit to before the cell runs. It is graded on the reasoning, not the guess, and being wrong here is the point.
  • Warmups: a one line blank or a two to four line exercise under the theory it practices, each with a four rung hint ladder behind it, where the last rung explains and still does not hand over the code.
  • An exercise that is broken when you open it.
  • A diagnose cell: code that runs, prints a confident and plausible answer, and is wrong. Something below it refuses the answer by computing the same thing a second way, so nothing is taken on trust.
  • A challenge that ends in a table and a sentence you write. The tutor grades the sentence, which means a green tick you earned for the wrong reason can be taken back.

No cell in this course passes in the state it ships. That is deliberate, and it is checked mechanically before the course is published.

The particular danger of this subject is an alarm that is true and useless. A fifth of a normal week's answers are wrong and that is not an incident. A deterministic judge does not make a ten row mean exact. A rule fitted on seven points fires one clean week in three. One alarm in the shift's window is a lucky draw, not detection. The topic that fell is two questions inside it. Doubling the rows does not halve the noise. A ten percent rise in latency is one column on one day. The retrieved document is a reference the system's own faults control. Paging on every alarm is paging nobody. Every diagnose cell is one of those, and every cross check is the second route that refuses it.

What you will be able to do

  • Read a production stream as rows with telemetry and measure what a normal week looks like, including its baseline rate of wrong answers.
  • Draw a seeded daily sample, judge it live, and measure the spread of a sample's mean.
  • Set a baseline, write a threshold rule and a cumulative one, and measure each rule's false alarm rate on a week with nothing in it.
  • Find a shift the mean hides and make it a fact by slicing to the questions it touched, with a row floor.
  • Size the sample from the row variance and read detection as a rate over redraws.
  • Tell a system change from a world change by which telemetry columns moved, and write a different runbook line for each.
  • Show why a judge referenced on the system's own source cannot see the world move and goes blind when the system breaks.
  • Write the monitor as a design with a reason beside every number, a cost, a false alarm rate and a runbook.

The lessons

1. Good on the test set. A normal week at 3.55 with a fifth of answers wrong. The census over three weeks: 3.55, 3.29, 2.86. Two events, planted so the method can be measured. A ten row monitor catches the outage on day 15 and the truth change on day 13, by luck.

2. Three weeks of traffic. Rows, topics, telemetry. The wrong share goes 19, 26, 38 percent. Week two's fall comes with no telemetry moving; week three's with every column.

3. The online judge. Ten seeded rows a day, judged live; the live scores match the recorded ones on nearly every row. The sample's mean sits 0.3 from the census on an average day and 0.8 on the worst, and a mean of ten spreads by 0.4.

4. The baseline and the alarm. Week one's sample gives 3.66 ± 0.46. The threshold fires on 13, 15, 16, 17, 20; CUSUM crosses on 13 and stays. The two sigma rule fitted on one clean week fires on another 33 times in 100.

5. The shift that hides. A quarter point in the mean is seen on 63 draws of 100 at ten rows. The two shifted questions fall from 4.32 to 2.41, five spreads of their own noise, while the rest holds at 3.43 against 3.42.

6. Sample size. Noise is 1.33 over root n. The outage is caught within two days on 56 draws at three rows, 82 at ten, 100 at twenty. The truth change would need about a hundred rows a day.

7. What else moved. Week three lights every telemetry column on every day from day 15; week two lights none beyond a normal week's one or two. On the shifted rows the telemetry is the same to the token.

8. The judge's reference. On the shifted questions the truth score falls 4.32 to 2.41 and the source score rises 3.14 to 3.44. The source referenced monitor alarms on no day and has no readings in week three.

9. Be the first to know. Twenty rows a day, two rules with measured false alarm rates, six telemetry columns, sliced questions with a floor, the truth as reference, and a runbook that reads coverage first. It names the world in week two and retrieval gone on day 15.

Requirements

  • Python 3.9 or later with ollama and pandas.
pip install ollama pandas
  • A running Ollama (https://ollama.com) with one model pulled once, used by lesson 3 only:
ollama pull qwen3:4b
  • EVAL 202 for the judge, EVAL 201 for labelled sets, PANDAS 101 for the tables.

The 630 rows and both judge scores are written into the setup cell, so nothing is downloaded and every lesson but the third runs without a model. Every number in the prose was produced by the cell above it on this machine; lesson 3's live scores may differ from the recorded ones on a handful of rows under another Ollama build, and every check there is a band rather than a typed answer.

What to read

For the alarm rules, search for control chart and CUSUM; the two rules in lesson 4 are the plain versions of Shewhart's and Page's, and the false alarm rate the lesson measures is what those texts call the average run length. For the subject, search for LLM observability with online evaluation and drift. EVAL 202 in this series is where the judge on every row was measured, EVAL 201 is where a labelled set is mined from a stream, and RED 301 is next: traffic that is hostile on purpose.