SLO lab
Triage and recover service-level objectives
When HTTP 200 fallbacks counted as good while citations empty — vanity availability, the system must degrade on purpose without widening blast radius.
1Learn the idea
Read
Classify and bound retries
Map failure classes for POST /v1/support/answer: retryable vs fatal vs needs-human. Retries need budgets, jitter, and idempotency rules aligned to bad events include malformed upstream + timeout fallbacks; auth 4xx excluded. The chapter’s signature failure — HTTP 200 fallbacks counted as good while citations empty — vanity availability — must take a deliberate branch, not a generic catch-all.
Read
Containment path
Implement the degrade/rollback/refuse behavior reliability owner gating releases on error budget needs when SLO-VANITY-200-18 repeats. Prefer scoped controls (one flag, one weight, one tenant, one secret version) over fleet-wide restarts. Preserve evidence; do not delete logs to “clean the demo.”
Read
Implementation artifact
- alert: SupportAnswerFastBurn
expr: slo:availability:burn_rate1h > 14.4 and slo:availability:burn_rate5m > 14.4
for: 2m
Read
Verify harm reduction
After containment, check slo:availability:ratio and burn_rate1h moves in the safe direction and watch for retry amplification. Write the stop condition that ends the incident response for this lab.
Read
Stage depth
Chaos note: inject only one fault class at a time and restore fixtures after. Watch for dual failures — dependency down and retry amplifier — which is how HTTP 200 fallbacks counted as good while citations empty — vanity availability becomes an outage. Customer communication templates (even if only for the drill) beat silence. If you queue deferred work, define poison-message handling. Budget documents should state the maximum extra spend allowed during retries. Close the loop by linking the containment action to a dashboard panel for slo:availability:ratio and burn_rate1h.
Read
Field notes for `slo-lab` / `failure-handling`
Draw a state diagram for degrade modes and put it in the repo as ASCII if needed. Cap concurrent retries across the process, not only per request. Ensure cancellation propagates to downstream HTTP clients. When failing closed, choose a user-visible message that does not leak internals. Practice the single command that flips the kill switch or weight to zero. After recovery, drain or inspect deferred work before declaring green. In this chapter the product is AI support-answer service with 99.5% availability + TTFT SLO, the human stakeholder is reliability owner gating releases on error budget, and the incident id you design against is SLO-VANITY-200-18. Re-state the oracle in your notes — 3% provider timeout for 20m trips fast-burn; monthly budget math matches calc-budget.py — and keep the invariant visible: bad events include malformed upstream + timeout fallbacks; auth 4xx excluded. Track slo:availability:ratio and burn_rate1h as the scoreboard. Surface under change control: POST /v1/support/answer. If you only have forty minutes, finish the fixture for HTTP 200 fallbacks counted as good while citations empty — vanity availability before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.
Go deeper
Before you start
Why this matters
Assume HTTP 200 fallbacks counted as good while citations empty — vanity availability is happening right now. Write the first safe action, the signal that confirms containment, and the action you will not take (infinite retry, broad restart, deleting evidence). Tie the plan to invariant: bad events include malformed upstream + timeout fallbacks; auth 4xx excluded.
In the wild
See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.