Page 3 of 8~245 min topic

Health check lab

Run the health checks baseline

One clean transaction through **GET /livez and GET /readyz** must match the oracle: vector pause → readyz 503, livez 200, no restart loop; process deadlock → livez fails → restart.

~30 min this pageHappy path

1Learn the idea

Read

Order the successful transaction

Code the narrow path that serves cluster operator during vector DB maintenance window: accept → authorize/normalize → call dependency → validate → record. Keep stages named so a trace can show which boundary passed. Success must emit evidence useful to restart_count == 0 during dep outage and ready_pods ≥ 1 when deps healthy, not only a 200 with prose. Predict the observable for GET /livez and GET /readyz before running: vector pause → readyz 503, livez 200, no restart loop; process deadlock → livez fails → restart.

Read

Run with fakes first

Drive the path with recording fakes or local stubs. Assert call order and arguments. Idempotency keys or stable ids should keep retries from duplicating costly work where the product requires it. Product under test remains Kubernetes liveness vs readiness for retrieval-and-generation API — resist adding unrelated features mid-path.

Read

Implementation artifact

curl -sf localhost:8080/livez && curl -sf localhost:8080/readyz

Read

Compare prediction to result

For Health check lab, paste the CLI/HTTP transcript beside your prediction for GET /livez and GET /readyz. If the oracle is unmet (vector pause → readyz 503, livez 200, no restart loop; process deadlock → livez fails → restart), stop and debug this page; do not compensate with prompt folktales. Re-run once after a clean process start to catch hidden global state that would invalidate PROBE-RESTART-STORM-8.

Read

Stage depth

Performance sketch: measure local p95 for the fake-backed path so later regressions are obvious. Keep concurrency modest until failure-handling proves limits. Log a single structured event per success with request id, revision, and the evidence field behind restart_count == 0 during dep outage and ready_pods ≥ 1 when deps healthy. Avoid hidden global caches in the happy path unless the lab is about caching — and even then key by tenant. If the path calls a model, pin model id in config and echo it in the response for auditability. Remember cluster operator during vector DB maintenance window experiences wall-clock time, not your debugger’s single-step comfort.

Read

Field notes for `health-check-lab` / `happy-path`

Prefer explicit function names over a single god-object handleRequest. Thread a correlation id from ingress to the last log line. When streaming, define what partial failure means before coding. Snapshot one successful response body in fixtures after redaction. If the path writes to a queue, assert message attributes in the fake. Stop adding retries on this page; that is the next concern. In this chapter the product is Kubernetes liveness vs readiness for retrieval-and-generation API, the human stakeholder is cluster operator during vector DB maintenance window, and the incident id you design against is PROBE-RESTART-STORM-8. Re-state the oracle in your notes — vector pause → readyz 503, livez 200, no restart loop; process deadlock → livez fails → restart — and keep the invariant visible: /livez only checks process; /readyz checks vector+model deps; wrong probe kills healthy pods. Track restart_count == 0 during dep outage and ready_pods ≥ 1 when deps healthy as the scoreboard. Surface under change control: GET /livez and GET /readyz. If you only have forty minutes, finish the fixture for readyz embedded in liveness → restart storm during dependency blip before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.

Go deeper

Before you start

Why this matters

Without calling production, order the steps a single success takes for cluster operator during vector DB maintenance window. Circle the first irreversible side effect. Your prediction should mention GET /livez and GET /readyz and the evidence field that proves vector pause → readyz 503, livez 200, no restart loop; process deadlock → livez fails → restart.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. Is call order asserted, not assumed?
2. Does success evidence support restart_count == 0 during dep outage and ready_pods ≥ 1 when deps healthy?
3. Did you compare prediction vs transcript?

All responses are required.