Capstone: research bot with citations
Set corpus trust and disclosure boundaries
Least privilege, negative probes, and a timed rollback beat a security essay about citation-first research assistant that maps every claim to evidence IDs.
1Learn the idea
Read
Least privilege for this lab
Separate runtime and operator roles for citation-first research assistant that maps every claim to evidence IDs. Runtime may only perform the narrow actions that analyst compiling a brief on bike-share policy from a fixed corpus needs; operators get audited break-glass with TTL. Encode a negative probe that denies the privilege trick related to model invents kb://minutes-2099 citation.
Read
Data and secret hygiene
Redact prompts/PII at collection. Secrets enter via a manager or workload identity — never source, fixtures, or exception strings. Incident CAP-RESEARCH-FAKE-ID-3 should be impossible if these controls hold. Output allowlists and schema checks stay in force on error paths.
Read
Implementation artifact
Read
Corpus is allowlisted paths only; reject http URLs in evidence_ids.
Read
Rollback drill
Rehearse the rollback or kill switch timed against a clock. Record actor, reason, prior revision/secret/flag, and verification query. Invariant reminder: claim without evidence → abstain or drop claim; no orphan sentences.
Read
Stage depth
Abuse cases unique to this lab include the privilege path implied by model invents kb://minutes-2099 citation. Prove a read-only role cannot mutate. Break-glass tokens expire; leftover tokens fail the drill. Dependency pin/digest story matters when images or models move under you. Document how to rotate the credential that citation-first research assistant that maps every claim to evidence IDs uses without a full outage window longer than your dual-run plan. Security evidence is part of ship, not an appendix nobody reads.
Read
Field notes for `capstone-research-bot` / `security-ops`
List network egress destinations and justify each. Ensure debug endpoints are off by default in the shipping config. Verify that error responses do not echo secrets or raw stack frames to clients. For multi-tenant paths, add a cross-tenant probe fixture. Time the rollback drill twice — once with the author, once with a peer. Store the drill transcript beside the threat notes for the incident id. In this chapter the product is citation-first research assistant that maps every claim to evidence IDs, the human stakeholder is analyst compiling a brief on bike-share policy from a fixed corpus, and the incident id you design against is CAP-RESEARCH-FAKE-ID-3. Re-state the oracle in your notes — brief on fixture corpus: ≥ 0.95 claims cited; fabricated ID rate 0 — and keep the invariant visible: claim without evidence → abstain or drop claim; no orphan sentences. Track claim_citation_coverage and fabricated_id_rate as the scoreboard. Surface under change control: POST /v1/research/brief. If you only have forty minutes, finish the fixture for model invents kb://minutes-2099 citation before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.
Go deeper
Before you start
Why this matters
Threat-model citation-first research assistant that maps every claim to evidence IDs in five minutes: who can change config, who can read secrets, what a malicious payload tries to do. Write one negative probe that must yield deny with zero side effects. Reference CAP-RESEARCH-FAKE-ID-3 as the story you refuse to repeat.
Capstone security gates are release blockers for citation-first research assistant that maps every claim to evidence IDs, not optional polish.
In the wild
See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.