Function calling in code
Ship and explain the order-status tool caller
Page 8 packages proved vs unproved evidence so another engineer can run, trust, or reject the order-status function-calling assistant.
1Learn the idea
Read
Assemble the ship record
A shippable lab artifact includes: how to run it, the metric result (tool-call validity rate; grounded status answers vs fabricated ones), the failure you can still reproduce (executing a hallucinated tool name, or answering status without a tool call), the security gate for allowing a write tool (refund/cancel) in the same registry, and a rollback note. The user decision it supports remains: let the model request a read-only get_order_status tool, then answer from the tool result.
Read
Freeze the evidence
npm run typecheck
npm test
Read
Run the credentialed smoke test only in an approved environment:
RUN_LIVE_EVALS=1 npm run test:live
Expected evidence: **A request for ord_123 produces one validated lookup and a reply based only on its result.**. Store this beside the fixture version so scores remain meaningful after content changes in `function-calling-code`.
Read
Explain limits without apology
State operating limits for the order-status tool caller in plain language: fixture size, offline vs live dependencies, and what would require a new eval set. Shipping function-calling-code is honest scoping, not maximal confidence language.
Read
Lab notebook: proved vs unproved
Fill this table in your notes for the order-status tool caller:
- Proved on
orders map + tool schema for get_order_status: … - Unproved beyond the fixture: …
- Metric that blocks release: tool-call validity rate; grounded status answers vs fabricated ones
- Failure still reproducible: executing a hallucinated tool name, or answering status without a tool call
- Security gate: allowing a write tool (refund/cancel) in the same registry
- Rollback: …
Ship the narrative only when the unproved list is honest. Reviewers trust narrow claims that support let the model request a read-only get_order_status tool, then answer from the tool result more than maximal language that collapses under the first production oddity.
Read
Worked judgment
Hand your ship note to a peer and ask them to recreate a proved/unproved ship note with rollback without watching you type. If they cannot, your evidence is still tribal knowledge. Tighten the run command and the metric line until a stranger can validate the order-status tool caller against orders map + tool schema for get_order_status.
Read
Why this stage matters for the order-status tool caller
At the mastery and shipping stage for function-calling-code, the job is narrower than finishing a product demo. You are creating one progressive evidence piece about orders map + tool schema for get_order_status that later pages inherit without redefining success. Keep that fixture small enough to inspect by hand, keep outputs copy-pasteable as text, and refuse to narrate this baseline as if it were a production SLA: assistant that answers without tools on the same prompts.
For this page specifically, success looks like a proved/unproved ship note with rollback while still centering the user decision to let the model request a read-only get_order_status tool, then answer from the tool result. If you cannot point to a file, command, or assertion that proves that for the order-status tool caller, stay on this page instead of advancing.
Glossary: tool · Glossary: structured output · Cheatsheet: production ops signals
Go deeper
Before you start
Why this matters
List two things this chapter proved on the fixture and two things it did not prove about the order-status tool caller. If you cannot name the gaps, you are not ready to ship the narrative—even if the code runs.
In the wild
See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.