Page 7 of 8~245 min topic

On-call lab

Secure on-call operations operations

Least privilege, negative probes, and a timed rollback beat a security essay about humane escalation path for production AI answer platform.

~25 min this pageSecurity and ops

1Learn the idea

Read

Least privilege for this lab

Separate runtime and operator roles for humane escalation path for production AI answer platform. Runtime may only perform the narrow actions that primary engineer receiving SEV2 after hours needs; operators get audited break-glass with TTL. Encode a negative probe that denies the privilege trick related to everyone paged for SEV3 noise → real SEV1 ignored.

Read

Data and secret hygiene

Redact prompts/PII at collection. Secrets enter via a manager or workload identity — never source, fixtures, or exception strings. Incident INC-OC-554 should be impossible if these controls hold. Output allowlists and schema checks stay in force on error paths.

Read

Implementation artifact

break_glass_ttl_minutes: 60
require_reason: true

Read

Rollback drill

Rehearse the rollback or kill switch timed against a clock. Record actor, reason, prior revision/secret/flag, and verification query. Invariant reminder: missed ACK in 15m escalates; SEV1 pages dual; handoff includes timeline + next action.

Read

Stage depth

Abuse cases unique to this lab include the privilege path implied by everyone paged for SEV3 noise → real SEV1 ignored. Prove a read-only role cannot mutate. Break-glass tokens expire; leftover tokens fail the drill. Dependency pin/digest story matters when images or models move under you. Document how to rotate the credential that humane escalation path for production AI answer platform uses without a full outage window longer than your dual-run plan. Security evidence is part of ship, not an appendix nobody reads.

Read

Field notes for `on-call-lab` / `security-ops`

List network egress destinations and justify each. Ensure debug endpoints are off by default in the shipping config. Verify that error responses do not echo secrets or raw stack frames to clients. For multi-tenant paths, add a cross-tenant probe fixture. Time the rollback drill twice — once with the author, once with a peer. Store the drill transcript beside the threat notes for the incident id. In this chapter the product is humane escalation path for production AI answer platform, the human stakeholder is primary engineer receiving SEV2 after hours, and the incident id you design against is INC-OC-554. Re-state the oracle in your notes — game day: primary misses ACK → secondary owns INC-OC-554 in 15m with access grant — and keep the invariant visible: missed ACK in 15m escalates; SEV1 pages dual; handoff includes timeline + next action. Track ack_minutes and pages_per_week as the scoreboard. Surface under change control: PagerDuty schedule ai-answer-primary. If you only have forty minutes, finish the fixture for everyone paged for SEV3 noise → real SEV1 ignored before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.

Go deeper

Before you start

Why this matters

Threat-model humane escalation path for production AI answer platform in five minutes: who can change config, who can read secrets, what a malicious payload tries to do. Write one negative probe that must yield deny with zero side effects. Reference INC-OC-554 as the story you refuse to repeat.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. Does deny imply zero side effects?
2. Are secrets absent from logs?
3. Is rollback evidenced with timestamps?

All responses are required.