Page 1 of 8~245 min topic

Runbook lab

Define the production target for runbooks

Ship a falsifiable slice of **executable queue-backlog runbook for embedding ingestion pipeline** — success is drill INC-INGEST-9021: lag 60→5 minutes in ≤ 25m following runbook v1.4, not a polished screenshot.

~25 min this pageLab goal

1Try it yourself

Decision drill

Runbook steps

Runbooks turn panic into ordered steps — page, rollback, document.

Runbook discipline69%

1/3PII in trace export — first action?

2Learn the idea

Read

Name the operable slice

This lab builds executable queue-backlog runbook for embedding ingestion pipeline. The human in the loop is secondary on-call clearing ingest.lag_minutes > 45. Scope is intentionally narrower than “make AI reliable”: you will prove one oracle — drill INC-INGEST-9021: lag 60→5 minutes in ≤ 25m following runbook v1.4 — and one invariant — every command is copy-pasteable; mutating steps need change ticket; verify after each step. Record non-goals in your notes so a later change cannot silently expand authority. The incident mnemonic for the chapter is INC-INGEST-9021; design as if that ticket is already written and you are filling evidence.

Read

Write the acceptance contract

Turn the oracle into a table: input fixture, expected observable, prohibited side effect, owner, latency/cost ceiling. Separate model taste from software correctness — transport, auth, parsing, and termination must be deterministic even when generated text varies. Primary metric family: drill_completion_minutes and lag_minutes. Averages without a denominator or revision label do not gate release. Fake external dependencies in unit tests; live calls wait until fakes pass.

Read

Implementation artifact

Read

Runbook: embedding ingest lag

Alert: ingest_lag_minutes > 45 for 10m Incident drill id: INC-INGEST-9021

Read

Freeze the first red test

Before implementation, encode a failing check that would have caught runbook says 'restart something' without naming deployment or verifying lag. That failure is the pedagogical north star for later pages: contracts reject it, happy path never performs it, validation asserts it, failure-handling contains it, observability detects it, security-ops prevents privilege tricks around it, and mastery replays it in a drill. Endpoint under study: ingest worker + queue metrics.

Read

Stage depth

Capacity note for planners: estimate peak demand on ingest worker + queue metrics and the cost ceiling for a failed retry storm. Write the abort conditions — unbounded spend, cross-tenant leakage, or inability to roll back — before you enjoy the first green test. Prefer synthetic fixtures shaped like production over anonymized production dumps you cannot share in class. When you are tempted to widen scope, re-read the oracle (drill INC-INGEST-9021: lag 60→5 minutes in ≤ 25m following runbook v1.4) and cut features that do not serve it. The teaching outcome is judgment under constraints: secondary on-call clearing ingest.lag_minutes > 45 gets a trustworthy control, not a kitchen-sink framework. Keep the language of release decisions: promote, hold, or roll back — never “see if it gets better.”

Read

Field notes for `runbook-lab` / `lab-goal`

Decide what will live in version control on day one: fixtures, contract markdown, and a failing test name. Write the cost ceiling as a hard number with currency and period. If the lab involves clusters, name the non-prod context you will use and forbid prod kubecontexts in scripts. Capture the baseline metric once before changing code so later gains are comparative. Refuse tools that hide the request path behind magic macros until the oracle is green on fakes. Your README section for this page should be five lines or fewer and still falsifiable. In this chapter the product is executable queue-backlog runbook for embedding ingestion pipeline, the human stakeholder is secondary on-call clearing ingest.lag_minutes > 45, and the incident id you design against is INC-INGEST-9021. Re-state the oracle in your notes — drill INC-INGEST-9021: lag 60→5 minutes in ≤ 25m following runbook v1.4 — and keep the invariant visible: every command is copy-pasteable; mutating steps need change ticket; verify after each step. Track drill_completion_minutes and lag_minutes as the scoreboard. Surface under change control: ingest worker + queue metrics. If you only have forty minutes, finish the fixture for runbook says 'restart something' without naming deployment or verifying lag before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.

Go deeper

Before you start

Why this matters

Write the single done-definition a reviewer would accept for Runbook lab (INC-INGEST-9021). Include the numeric gate hidden in this oracle: drill INC-INGEST-9021: lag 60→5 minutes in ≤ 25m following runbook v1.4. Then name the fake success you refuse: a demo that ignores runbook says 'restart something' without naming deployment or verifying lag. Keep the sentence beside your editor; every later page should make this sentence easier to prove.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. Is the oracle (drill INC-INGEST-9021: lag 60→5 minutes in ≤ 25m following runbook v1.4) falsifiable from a fixture?
2. Is the invariant (every command is copy-pasteable; mutating steps need change ticket; verify after each step) stated without hand-waving?
3. Does the contract name INC-INGEST-9021 as a risk you design against?

All responses are required.