Multimodal API lab
Frame the typed receipt-vision API experiment
Ship a falsifiable slice of **typed receipt-vision API that extracts total + merchant with uncertainty** — success is sharp receipt → total=24.50 merchant=Cafe Nora confidence≥0.7; blurry → abstain, not a polished screenshot.
1Try it yourself
Playground
Vision API wiring
Image in → text out — same chat API, multimodal message parts.
- Base64 or URL-encode image
- User message with image + text parts
- POST chat/completions with vision model
- Parse text answer — verify before act
2Learn the idea
Read
Name the operable slice
This lab builds typed receipt-vision API that extracts total + merchant with uncertainty. The human in the loop is expense bot that must refuse blurry or non-receipt images. Scope is intentionally narrower than “make AI reliable”: you will prove one oracle — sharp receipt → total=24.50 merchant=Cafe Nora confidence≥0.7; blurry → abstain — and one invariant — image ≤ 4 MiB, MIME allowlisted, model output schema-validated before DB write. Record non-goals in your notes so a later change cannot silently expand authority. The incident mnemonic for the chapter is VISION-MEME-77; design as if that ticket is already written and you are filling evidence.
Read
Write the acceptance contract
Turn the oracle into a table: input fixture, expected observable, prohibited side effect, owner, latency/cost ceiling. Separate model taste from software correctness — transport, auth, parsing, and termination must be deterministic even when generated text varies. Primary metric family: extraction_precision and abstain_on_non_receipt ≥ 0.95. Averages without a denominator or revision label do not gate release. Fake external dependencies in unit tests; live calls wait until fakes pass.
Read
Implementation artifact
export const acceptance = {
maxBytes: 4 * 1024 * 1024,
mime: ["image/jpeg", "image/png", "image/webp"],
oracle: "Cafe Nora / 24.50 or abstain",
};
Read
Freeze the first red test
Before implementation, encode a failing check that would have caught screenshot of a meme parsed as a $9,999 expense. That failure is the pedagogical north star for later pages: contracts reject it, happy path never performs it, validation asserts it, failure-handling contains it, observability detects it, security-ops prevents privilege tricks around it, and mastery replays it in a drill. Endpoint under study: POST /v1/receipts/extract.
Read
Stage depth
Capacity note for planners: estimate peak demand on POST /v1/receipts/extract and the cost ceiling for a failed retry storm. Write the abort conditions — unbounded spend, cross-tenant leakage, or inability to roll back — before you enjoy the first green test. Prefer synthetic fixtures shaped like production over anonymized production dumps you cannot share in class. When you are tempted to widen scope, re-read the oracle (sharp receipt → total=24.50 merchant=Cafe Nora confidence≥0.7; blurry → abstain) and cut features that do not serve it. The teaching outcome is judgment under constraints: expense bot that must refuse blurry or non-receipt images gets a trustworthy control, not a kitchen-sink framework. Keep the language of release decisions: promote, hold, or roll back — never “see if it gets better.”
Read
Field notes for `multimodal-api-lab` / `lab-goal`
Decide what will live in version control on day one: fixtures, contract markdown, and a failing test name. Write the cost ceiling as a hard number with currency and period. If the lab involves clusters, name the non-prod context you will use and forbid prod kubecontexts in scripts. Capture the baseline metric once before changing code so later gains are comparative. Refuse tools that hide the request path behind magic macros until the oracle is green on fakes. Your README section for this page should be five lines or fewer and still falsifiable. In this chapter the product is typed receipt-vision API that extracts total + merchant with uncertainty, the human stakeholder is expense bot that must refuse blurry or non-receipt images, and the incident id you design against is VISION-MEME-77. Re-state the oracle in your notes — sharp receipt → total=24.50 merchant=Cafe Nora confidence≥0.7; blurry → abstain — and keep the invariant visible: image ≤ 4 MiB, MIME allowlisted, model output schema-validated before DB write. Track extraction_precision and abstain_on_non_receipt ≥ 0.95 as the scoreboard. Surface under change control: POST /v1/receipts/extract. If you only have forty minutes, finish the fixture for screenshot of a meme parsed as a $9,999 expense before polishing UI. Promotion language stays ternary: promote, hold, or roll back based on evidence, not hope.
Go deeper
Before you start
Why this matters
Write the single done-definition a reviewer would accept for Multimodal API lab (VISION-MEME-77). Include the numeric gate hidden in this oracle: sharp receipt → total=24.50 merchant=Cafe Nora confidence≥0.7; blurry → abstain. Then name the fake success you refuse: a demo that ignores screenshot of a meme parsed as a $9,999 expense. Keep the sentence beside your editor; every later page should make this sentence easier to prove.
In the wild
See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.