Page 2 of 8~112 min topic

Eval metrics lab

Set up interfaces and contracts

Types and env contracts make **offline+online eval harness for grounded support answers** fail closed before any provider or cluster call.

~14 min this pageSetup

1Learn the idea

Read

Define trusted borders

Implement types or schemas around python -m eval.run --gold v12 so illegal states are unrepresentable at the boundary. For offline+online eval harness for grounded support answers, trusted inputs come from sessions, signatures, pinned digests, or workload identity — not from free-form model text. ML engineer gating a prompt change before Friday release should be able to read the contract and know which fields are optional, which are enumerated, and which abort the request. Constructors and boot paths must not perform provider side effects; injection points keep tests honest.

Read

Environment and secrets contract

Document required env vars in .env.example with placeholders only. Rotation story belongs later, but the contract already forbids printing secrets and forbids defaulting to fail-open when a dependency is missing. Invariant to encode in types/tests: release needs groundedness ≥ 0.88 and latency p95 ≤ 2.0s on fixed gold set v12. If a config number lacks units, fix the name (timeout_ms, rpm_hard) before writing logic.

Read

Implementation artifact

@dataclass
class CaseResult:
    id: str
    grounded: bool
    citation_ok: bool
    latency_ms: int

Read

Fixture kit

Create fixtures for the golden path and for EVAL-LEAK-308. Name files after the behavior (429-retry-after.json, cross-tenant.json, empty-citation.json) rather than test1. Each fixture carries expected status/code. This kit is the shared language for validation and failure pages.

Read

Stage depth

Compatibility promise: additive fields may appear only if readers ignore unknowns safely; breaking changes bump a version visible on the wire. For AI payloads, size limits arrive before JSON parse when hostile blobs are a risk. Document how clock skew, idempotency keys, and tracing headers travel through offline+online eval harness for grounded support answers. If you use feature flags later, the contract already states that flags are not authorization. Link each config knob to a unit and a failure mode (“0 means disabled” vs “0 means divide-by-zero”). A peer reviewing the PR should find EVAL-LEAK-308 named in a comment on the adversarial fixture.

Read

Field notes for `eval-metrics-lab` / `setup-and-contract`

Generate OpenAPI or a typed client only after the hand schema is stable for one fixture round-trip. Record how errors look on the wire — problem+json, envelope, or bare status — and stick to one. Clock sources must be injectable for skew tests. If webhooks appear later, document signature header names now even as TODOs. Keep sample payloads UTF-8 and free of real emails. Add a makefile or npm script that validates schemas without network. In this chapter the product is offline+online eval harness for grounded support answers, the human stakeholder is ML engineer gating a prompt change before Friday release, and the incident id you design against is EVAL-LEAK-308. Re-state the oracle in your notes — candidate prompt beats baseline on groundedness by ≥ 0.03 without latency regression > 10% — and keep the invariant visible: release needs groundedness ≥ 0.88 and latency p95 ≤ 2.0s on fixed gold set v12. Track groundedness, citation_precision, p95_latency_ms as the scoreboard. Surface under change control: python -m eval.run --gold v12. If you only have forty minutes, finish the fixture for eval set leaked into few-shot examples — scores look perfect, prod drops before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.

Go deeper

Before you start

Why this matters

For Eval metrics lab, sketch the request and response shapes that cross python -m eval.run --gold v12 without naming a framework. Mark which fields are trusted (session, signatures, digests) versus untrusted (user text, model JSON, webhook bodies). If a field can change authorization, it does not belong in model output. Predict one 422/401 you will assert before coding adapters for offline+online eval harness for grounded support answers.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. Can a newcomer list trusted vs untrusted fields?
2. Does boot avoid external side effects?
3. Are fixtures named after EVAL-LEAK-308-class failures?

All responses are required.