Eval metrics lab
Set up interfaces and contracts
Types and env contracts make **offline+online eval harness for grounded support answers** fail closed before any provider or cluster call.
1Learn the idea
Read
Define trusted borders
Implement types or schemas around python -m eval.run --gold v12 so illegal states are unrepresentable at the boundary. For offline+online eval harness for grounded support answers, trusted inputs come from sessions, signatures, pinned digests, or workload identity — not from free-form model text. ML engineer gating a prompt change before Friday release should be able to read the contract and know which fields are optional, which are enumerated, and which abort the request. Constructors and boot paths must not perform provider side effects; injection points keep tests honest.
Read
Environment and secrets contract
Document required env vars in .env.example with placeholders only. Rotation story belongs later, but the contract already forbids printing secrets and forbids defaulting to fail-open when a dependency is missing. Invariant to encode in types/tests: release needs groundedness ≥ 0.88 and latency p95 ≤ 2.0s on fixed gold set v12. If a config number lacks units, fix the name (timeout_ms, rpm_hard) before writing logic.
Read
Implementation artifact
@dataclass
class CaseResult:
id: str
grounded: bool
citation_ok: bool
latency_ms: int
Read
Fixture kit
Create fixtures for the golden path and for EVAL-LEAK-308. Name files after the behavior (429-retry-after.json, cross-tenant.json, empty-citation.json) rather than test1. Each fixture carries expected status/code. This kit is the shared language for validation and failure pages.
Read
Stage depth
Compatibility promise: additive fields may appear only if readers ignore unknowns safely; breaking changes bump a version visible on the wire. For AI payloads, size limits arrive before JSON parse when hostile blobs are a risk. Document how clock skew, idempotency keys, and tracing headers travel through offline+online eval harness for grounded support answers. If you use feature flags later, the contract already states that flags are not authorization. Link each config knob to a unit and a failure mode (“0 means disabled” vs “0 means divide-by-zero”). A peer reviewing the PR should find EVAL-LEAK-308 named in a comment on the adversarial fixture.
Read
Field notes for `eval-metrics-lab` / `setup-and-contract`
Generate OpenAPI or a typed client only after the hand schema is stable for one fixture round-trip. Record how errors look on the wire — problem+json, envelope, or bare status — and stick to one. Clock sources must be injectable for skew tests. If webhooks appear later, document signature header names now even as TODOs. Keep sample payloads UTF-8 and free of real emails. Add a makefile or npm script that validates schemas without network. In this chapter the product is offline+online eval harness for grounded support answers, the human stakeholder is ML engineer gating a prompt change before Friday release, and the incident id you design against is EVAL-LEAK-308. Re-state the oracle in your notes — candidate prompt beats baseline on groundedness by ≥ 0.03 without latency regression > 10% — and keep the invariant visible: release needs groundedness ≥ 0.88 and latency p95 ≤ 2.0s on fixed gold set v12. Track groundedness, citation_precision, p95_latency_ms as the scoreboard. Surface under change control: python -m eval.run --gold v12. If you only have forty minutes, finish the fixture for eval set leaked into few-shot examples — scores look perfect, prod drops before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.
Go deeper
Before you start
Why this matters
For Eval metrics lab, sketch the request and response shapes that cross python -m eval.run --gold v12 without naming a framework. Mark which fields are trusted (session, signatures, digests) versus untrusted (user text, model JSON, webhook bodies). If a field can change authorization, it does not belong in model output. Predict one 422/401 you will assert before coding adapters for offline+online eval harness for grounded support answers.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.