Page 3 of 8~112 min topic

A/B test lab

Implement one traceable happy path

One clean transaction through **POST /v1/support/answer** must match the oracle: after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget.

~14 min this pageVertical slice

1Try it yourself

Decision drill

A/B test lab

Split traffic and measure quality — not every change needs an experiment, but behavior shifts do.

Experiment rigor71%

1/3New system prompt — 50/50 traffic.

2Learn the idea

Read

Order the successful transaction

Code the narrow path that serves experiment owner measuring groundedness lift with SRM checks: accept → authorize/normalize → call dependency → validate → record. Keep stages named so a trace can show which boundary passed. Success must emit evidence useful to groundedness_lift_pp and srm_pvalue, not only a 200 with prose. Predict the observable for POST /v1/support/answer before running: after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget.

Read

Run with fakes first

Drive the path with recording fakes or local stubs. Assert call order and arguments. Idempotency keys or stable ids should keep retries from duplicating costly work where the product requires it. Product under test remains support-answer service comparing retriever_v1 vs reranking retriever_v2 — resist adding unrelated features mid-path.

Read

Implementation artifact

arm = assign("user_42")
answer = answer_with(retrievers[arm], q)
log_exposure(user_id="user_42", arm=arm, exp=EXP["key"])

Read

Compare prediction to result

For A/B test lab, paste the CLI/HTTP transcript beside your prediction for POST /v1/support/answer. If the oracle is unmet (after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget), stop and debug this page; do not compensate with prompt folktales. Re-run once after a clean process start to catch hidden global state that would invalidate AB-SRM-FAIL-27.

Read

Stage depth

Performance sketch: measure local p95 for the fake-backed path so later regressions are obvious. Keep concurrency modest until failure-handling proves limits. Log a single structured event per success with request id, revision, and the evidence field behind groundedness_lift_pp and srm_pvalue. Avoid hidden global caches in the happy path unless the lab is about caching — and even then key by tenant. If the path calls a model, pin model id in config and echo it in the response for auditability. Remember experiment owner measuring groundedness lift with SRM checks experiences wall-clock time, not your debugger’s single-step comfort.

Read

Field notes for `ab-test-lab` / `happy-path`

Prefer explicit function names over a single god-object handleRequest. Thread a correlation id from ingress to the last log line. When streaming, define what partial failure means before coding. Snapshot one successful response body in fixtures after redaction. If the path writes to a queue, assert message attributes in the fake. Stop adding retries on this page; that is the next concern. In this chapter the product is support-answer service comparing retriever_v1 vs reranking retriever_v2, the human stakeholder is experiment owner measuring groundedness lift with SRM checks, and the incident id you design against is AB-SRM-FAIL-27. Re-state the oracle in your notes — after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget — and keep the invariant visible: assignment sticky by user_id; analysis gated on sample ratio mismatch (SRM) p>0.001 fail. Track groundedness_lift_pp and srm_pvalue as the scoreboard. Surface under change control: POST /v1/support/answer. If you only have forty minutes, finish the fixture for reassignment every request → users flicker; metrics uninterpretable before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.

Go deeper

Before you start

Why this matters

Without calling production, order the steps a single success takes for experiment owner measuring groundedness lift with SRM checks. Circle the first irreversible side effect. Your prediction should mention POST /v1/support/answer and the evidence field that proves after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. Is call order asserted, not assumed?
2. Does success evidence support groundedness_lift_pp and srm_pvalue?
3. Did you compare prediction vs transcript?

All responses are required.