Multi-agent in code
Frame the research-critique coordinator experiment
Page 1 sets a falsifiable claim for the research+critique coordinator before any implementation work begins.
1Try it yourself
Playground
Multi-agent routing
Orchestrator plans and merges — workers do bounded subtasks with step caps.
Gather docs from 3 sources
2Learn the idea
Read
Name the deliverable and claim
Success is not “I followed the tutorial.” Success is producing evidence that: critique can veto an unsupported research draft; coordinator never exceeds worker budgets. The accepted input is narrow on purpose: user brief, research worker, critique worker, merge policy. That narrowness is what lets you inspect every field and prevents a toy demo from being narrated as a production system.
Record the baseline you must beat: single agent without critique pass. If the finished artifact cannot beat that baseline on the fixture below, stop and revise the claim before writing more code.
Read
Inventory the fixture
export const acceptance = {
project: "a coordinator that delegates research and critique to two bounded specialist workers",
invariant: "workers exchange typed artifacts rather than unconstrained chat transcripts",
expected: "Research returns two sourced claims; critique rejects one unsupported claim; synthesis keeps one.",
} as const;
Expected evidence: Research returns two sourced claims; critique rejects one unsupported claim; synthesis keeps one.. Treat the printout as a claim about this fixture, not as proof that the toolchain merely started.
Read
Spot misleading success early
For the research+critique coordinator, a decorative win often looks like a clean run that never checks veto rate; unsupported claim count after merge; total step budget. Write the metric down now so later pages cannot redefine success after the fact. Also note the operational threat you will eventually gate on: critique worker receiving secrets from research scratchpad logs.
Read
Lab notebook: claim before code
For multi-agent-in-code, write the claim on a sticky note in this exact shape: “Given user brief, research worker, critique worker, merge policy, the research-critique coordinator will …”. Fill the ellipsis with the observable part of: critique can veto an unsupported research draft; coordinator never exceeds worker budgets. Tape the baseline beside it: single agent without critique pass. If someone later replaces your metric with a vibe check, the sticky note is how you push back.
Also sketch the one-sentence user story: a person uses this output to delegate research and critique to two bounded workers, then merge under a coordinator policy. If that sentence needs a dashboard, a model zoo, or five services, the lab scope is too wide—shrink the fixture (two workers with separate step budgets) until the story fits on one screen.
Read
Worked judgment
Decide now whether live network calls are allowed on page 1. For this lab they usually are not; inventory and contracts should run offline against two workers with separate step budgets. Note the metric you will eventually require (veto rate; unsupported claim count after merge; total step budget) so page 4 cannot invent a softer target. The characteristic failure to keep in mind is workers sharing mutable state, or coordinator ignoring critique veto.
Read
Why this stage matters for the research-critique coordinator
At the experiment brief stage for multi-agent-in-code, the job is narrower than finishing a product demo. You are creating one progressive evidence piece about two workers with separate step budgets that later pages inherit without redefining success. Keep that fixture small enough to inspect by hand, keep outputs copy-pasteable as text, and refuse to narrate this baseline as if it were a production SLA: single agent without critique pass.
For this page specifically, success looks like a falsifiable claim and baseline written before coding while still centering the user decision to delegate research and critique to two bounded workers, then merge under a coordinator policy. If you cannot point to a file, command, or assertion that proves that for the research-critique coordinator, stay on this page instead of advancing.
Glossary: tool · Glossary: structured output · Cheatsheet: production ops signals
Go deeper
Before you start
Why this matters
On paper, write the user decision this lab supports: delegate research and critique to two bounded workers, then merge under a coordinator policy. Then write one sentence naming what could look successful while actually being wrong for this claim—focus on workers sharing mutable state, or coordinator ignoring critique veto. Keep both sentences beside the fixture inventory you run next.
In the wild
See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.