Page 3 of 8~112 min topic

Guardrails in code

Build the first working pre/post guardrail pipeline

Page 3 implements the shortest complete path for the pre/post output guardrail pipeline with inspectable intermediate values.

~14 min this pageImplementation

1Learn the idea

Read

Implement the minimal working path

Build only what the claim requires: blocked inputs never call the model; unsafe outputs are refused with a stable code. Prefer boring, deterministic code over frameworks you cannot yet explain. Run the path twice; identical output on this fixture is a feature, not a lack of creativity.

Read

Run the working path

def run_case(case, *, release, trace_id):
    # Adapters are injected in production; the orchestration owns policy.
    decision = guard(request, allowed_tools={"lookup_order"}, bind_identity=True)
    reasons = evaluate(answer if "answer" in locals() else locals().get("result", locals().get("decision", locals().get("job", locals().get("choice")))))
    return {"passed": not reasons, "reasons": reasons,
            "release": release, "trace_id": trace_id}

result = run_case({"user_text":"Track order A12","session_customer":"cust_7","tool":"lookup_order","args":{"order_id":"A12","customer_id":"cust_99"}}, release="candidate", trace_id="trace-42")
assert result["trace_id"] == "trace-42"

Expected evidence: a keyword filter passes harmless wording but the model emits refund_order with an attacker-controlled customer_id. Read each printed intermediate as part of the argument that the path works—not as decoration.

Read

Trace one input end to end

Narrate the journey from raw input to result for a single example from policy with PII and self-harm categories + sample prompts. If you cannot name an intermediate, the implementation is still too opaque for this lab. Only after this path is solid should you generalize data sources or UI.

Read

Lab notebook: intermediates worth printing

While implementing the pre/post guardrail pipeline, print or log at least three intermediates that map to the claim (blocked inputs never call the model; unsafe outputs are refused with a stable code). Good intermediates are values a teammate could recompute with a calculator or diff. Bad intermediates are framework traces you cannot explain.

Re-run with policy with PII and self-harm categories + sample prompts twice. If the second run differs, either the path is nondeterministic (document the seed) or you have hidden global state—both are lab bugs until named.

Read

Worked judgment

Stop adding features once the path supports block unsafe prompts and strip or refuse unsafe completions before they reach users. Extra UI, extra tools, or extra models belong in later chapters. The mastery bar for this page is simply: a deterministic end-to-end path with intermediates.

Read

Why this stage matters for the pre/post guardrail pipeline

At the implementation stage for guardrails-in-code, the job is narrower than finishing a product demo. You are creating one progressive evidence piece about policy with PII and self-harm categories + sample prompts that later pages inherit without redefining success. Keep that fixture small enough to inspect by hand, keep outputs copy-pasteable as text, and refuse to narrate this baseline as if it were a production SLA: model call with no guards on the same fixtures.

For this page specifically, success looks like a deterministic path with printed intermediates while still centering the user decision to block unsafe prompts and strip or refuse unsafe completions before they reach users. If you cannot point to a file, command, or assertion that proves that for the pre/post guardrail pipeline, stay on this page instead of advancing.

Cheatsheet: prompt injection defense

Previous · Next

Go deeper

Before you start

Why this matters

Without running code, predict the final output for fixture policy with PII and self-harm categories + sample prompts. Name one intermediate value that would prove the prediction. Then answer: what could look successful while actually being wrong at this stage for the pre/post guardrail pipeline?

In the wild

See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. Can you narrate every intermediate value?
2. Is the fixture deterministic and independently inspectable?
3. Did you avoid framework behavior you cannot explain yet?
4. Does the output still support the decision: block unsafe prompts and strip or refuse unsafe completions before they reach users?

All responses are required.