A/B test lab
Prove behavior with deterministic tests
Executable checks prove assignment sticky by user_id; analysis gated on sample ratio mismatch (SRM) p>0.001 fail on fixtures — including the known misshape behind AB-SRM-FAIL-27.
1Try it yourself
Decision drill
A/B test lab
Split traffic and measure quality — not every change needs an experiment, but behavior shifts do.
1/3New system prompt — 50/50 traffic.
2Learn the idea
Read
Schema and policy checks
Add executable validation at the trust boundaries of support-answer service comparing retriever_v1 vs reranking retriever_v2. Reject unknown fields where they matter, bound string sizes, and coerce only after auth/signature checks when raw bytes are security-relevant. Invariant under test: assignment sticky by user_id; analysis gated on sample ratio mismatch (SRM) p>0.001 fail. A TypeScript type or Python annotation is not runtime validation — pair them with parsers.
Read
Golden and adversarial fixtures
Automate the fixtures from setup, including a recreation of AB-SRM-FAIL-27. Assert both the visible error and the absence of side effects (no provider call, no queue write, no flag flip). Where metrics matter, assert label enums stay bounded.
Read
Implementation artifact
srm_p = srm_test(exposures, expected=[0.5, 0.5])
assert srm_p > 0.001
Read
Gate semantics
Document which failures are client mistakes (4xx) versus operator/config mistakes (5xx/503). Oracle still stands: after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget. Validation should make accidental “success with empty body” impossible for experiment owner measuring groundedness lift with SRM checks.
Read
Stage depth
Property ideas: shuffled field order, Unicode edges, maximum-length strings, and replayed timestamps. Where money, identity, or citations matter, assertion messages should cite the field name. Do not snapshot entire provider payloads in tests; assert semantically. If validation fails open “to keep the demo working,” you have inverted the lab. Tie at least one CI job to the AB-SRM-FAIL-27 fixture so main cannot regress silently. Re-read assignment sticky by user_id; analysis gated on sample ratio mismatch (SRM) p>0.001 fail after each new parser — convenience helpers love to bypass it.
Read
Field notes for `ab-test-lab` / `validation`
Table-drive status codes and error codes so reviewers see coverage at a glance. Include a Unicode normalization case if user text is accepted. Verify that oversized bodies fail before CPU-heavy work. Where digests or versions are pinned, assert mismatch behavior. Keep golden files small enough to read in review. CI should fail on skipped tests that mark the incident fixture as xfail without a ticket link. In this chapter the product is support-answer service comparing retriever_v1 vs reranking retriever_v2, the human stakeholder is experiment owner measuring groundedness lift with SRM checks, and the incident id you design against is AB-SRM-FAIL-27. Re-state the oracle in your notes — after 20k sessions, v2 groundedness +2.1pp, p95 latency +80ms within budget — and keep the invariant visible: assignment sticky by user_id; analysis gated on sample ratio mismatch (SRM) p>0.001 fail. Track groundedness_lift_pp and srm_pvalue as the scoreboard. Surface under change control: POST /v1/support/answer. If you only have forty minutes, finish the fixture for reassignment every request → users flicker; metrics uninterpretable before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.
Go deeper
Before you start
Why this matters
List three fixtures: one golden success, one schema/auth reject, and one regression for AB-SRM-FAIL-27. For each, write the exact assertion (status, code, metric, or citation) that must turn red if broken.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.