Page 3 of 8~104 min topic

Prompt injection & AI security

Learn the controls and knobs

Each prompt injection control is a hypothesis about a metric under a workload—not a synonym for quality on the tool-using support agent.

~13 min this pageControls

1Learn the idea

Read

Control map

See it

Why fluent answers can still be wrong
01Predict ≠ lookupSounds like an answer
02Web is messyFacts + fanfic mix
03No embarrassmentCan sound sure
04Prompt trapAsked to invent detail

Confidence is a tone — verify before you act

Primary knobs for prompt injection: tool allowlists, argument schemas, human approval gates, content delimiters, retrieval trust levels, output filters.

Write a sheet for the tool-using support agent with columns: control, current value, predicted benefit, predicted cost, rollback trigger. Fill it using this topic’s real tension: Blocking suspicious phrases is simple but produces false positives and misses paraphrases. Giving an agent broad tools increases usefulness and blast radius together. Human approval reduces autonomous speed but is appropriate for payments, deletion, disclosure, and external messages.

Change one prompt injection family at a time. If you move two knobs and the tool-using support agent improves, you learned a cocktail, not a cause—and you cannot roll back surgically.

Read

Product exposure

End users of the tool-using support agent should see only safe dials related to prompt injection. Infrastructure limits, private prompts, and policy thresholds stay server-owned. A user-facing control that bypasses those limits is a vulnerability dressed as UX for prompt injection.

Read

Make it operational

Publish the prompt injection control sheet next to the tool-using support agent runbook. On-call should see which knob moved in the last deploy without reading chat archaeology. Unknown prompt injection knobs are unowned knobs.

Also pin one numeric memory from this prompt injection chapter: risk ≈ probability of successful injection × impact of available capability; reducing tool privilege cuts impact even when detection is imperfect That number is not decoration; it is a template for how claims about prompt injection on the tool-using support agent should look in design docs. Scoped specifically to prompt injection / tool-using support agent / controls-and-knobs.

Read

Common mix-ups

People confuse prompt injection with neighboring buzzwords when debugging the tool-using support agent. Before changing prompts, ask whether the broken stage was evidence gathering, the prompt injection judgment itself, validation, or the product action. Fixing the wrong stage creates folklore (“we tried prompt injection and it failed”) that blocks the next team on the tool-using support agent. Scoped specifically to prompt injection / tool-using support agent / controls-and-knobs.

Read

Rehearsal (prompt-injection/controls-and-knobs)

Write a five-line artifact for this page: goal, inputs, check, owner, stop rule. Invent one fluent failure that the check would catch. Keep details specific to prompt injection rather than generic AI advice.

Read

Rehearsal (prompt-injection/controls-and-knobs)

Write a five-line artifact for this page: goal, inputs, check, owner, stop rule. Invent one fluent failure that the check would catch. Keep details specific to prompt injection rather than generic AI advice.

Go deeper

Before you start

Why this matters

From [tool allowlists, argument schemas, human approval gates, content delimiters, retrieval trust levels, output filters], pick one control for prompt injection on the tool-using support agent. Predict which metric rises and which cost rises if you increase it.

In the wild

See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. What is one idea from this page you would apply, and what evidence would you check?

All responses are required.