Guardrails in code
Set release boundaries for the pre/post guardrail pipeline
Page 7 defines what the pre/post output guardrail pipeline must refuse before release—security here is not a pasted happy path.
1Learn the idea
Read
Threats for this artifact only
Operational risks for the pre/post output guardrail pipeline center on logging blocked prompts that contain secrets into an unredacted sink, plus the earlier failure mode (post-guard only that still bills the model for blocked intents, or regex gaps on obfuscation). Safety lives in executable gates, allowlists, redaction, and a named owner—not in a warning paragraph under an unsafe function.
Read
Run the release gate
DENIED={'api_key','authorization','raw_prompt'}
def public_log(d):
return {k:('[REDACTED]' if k in DENIED else v) for k,v in d.items()}
print(public_log({'rule':'pii','api_key':'sk','ok':True}))
Expected evidence: redacted guard log. A failed assertion means stop, investigate, and do not publish the pre/post guardrail pipeline.
Read
Owner, retention, rollback
Name who can disable the feature, what data is retained, and how to roll back to the last known good artifact. Pin the reviewed configuration (versions, thresholds, allowlists) so “what shipped” is reconstructable for guardrails-in-code.
Read
Lab notebook: release blocker
Write the release blocker as a predicate, not a feeling: “Do not ship the pre/post guardrail pipeline if logging blocked prompts that contain secrets into an unredacted sink.” Pair it with a passing control that shows the reviewed configuration still works for policy with PII and self-harm categories + sample prompts. Name an owner and a rollback handle (git tag, docs_version, previous image).
Security pages must not paste the happy-path demo. If your gate code looks like the implementation page, replace it with a deny/allow check aimed at logging blocked prompts that contain secrets into an unredacted sink.
Read
Worked judgment
State the data retention rule in one line (what is stored, for how long, who can read it). Then state the kill switch (env flag, config pin, or feature owner). The pre/post guardrail pipeline is not shippable without both, even when block rate, false-block samples, escape rate on red-team set looks healthy.
Read
Why this stage matters for the pre/post guardrail pipeline
At the safety and operations stage for guardrails-in-code, the job is narrower than finishing a product demo. You are creating one progressive evidence piece about policy with PII and self-harm categories + sample prompts that later pages inherit without redefining success. Keep that fixture small enough to inspect by hand, keep outputs copy-pasteable as text, and refuse to narrate this baseline as if it were a production SLA: model call with no guards on the same fixtures.
For this page specifically, success looks like an executable deny gate for the lab-specific threat while still centering the user decision to block unsafe prompts and strip or refuse unsafe completions before they reach users. If you cannot point to a file, command, or assertion that proves that for the pre/post guardrail pipeline, stay on this page instead of advancing.
Go deeper
Before you start
Why this matters
Write an attack or unsafe misuse specific to this lab: logging blocked prompts that contain secrets into an unredacted sink. Predict whether your current code blocks it. Then run the gate below and compare.
In the wild
See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.