Page 2 of 8~116 min topic

Why AI makes mistakes

A map of AI mistakes

Not every wrong output is a hallucination. Classifying the failure helps you choose a better repair than simply asking the model to try again.

~15 min this pageFailure-mode taxonomy

1Try it yourself

Simulation game

Hallucination hunt

Stamp each claim: Trap or Trust. Confident voice ≠ true.

Quiz show

SOUNDS SURE

Sydney is the capital of Australia.

2Learn the idea

Read

The core idea

See it

Why fluent answers can still be wrong
01Predict ≠ lookupSounds like an answer
02Web is messyFacts + fanfic mix
03No embarrassmentCan sound sure
04Prompt trapAsked to invent detail

Confidence is a tone — verify before you act

Not every wrong output is a hallucination. Classifying the failure helps you choose a better repair than simply asking the model to try again.

Read

A practical lens

Use this three-part method:

  1. Label fabrication, staleness, ambiguity, calculation, and retrieval failures separately. Write down what this means in the scenario, what evidence would show it was done, and who owns the decision.
  2. Locate the stage where the error entered. Write down what this means in the scenario, what evidence would show it was done, and who owns the decision.
  3. Apply a repair aimed at that stage. Write down what this means in the scenario, what evidence would show it was done, and who owns the decision.

Read

Worked example

Walk through A project brief contains an invented statistic, an outdated product name, and a correct fact attached to the wrong source.. Label three moments where “A map of AI mistakes” changes what you trust: (1) the first fluent answer, (2) the first missing source or permission, and (3) the decision a human must own. Write the before/after task so the model only does the slice that evidence supports. Keep one sentence that states how this page’s idea differs from a generic “AI is smart/dumb” score.

Read

Common traps and better moves

  • Using hallucination as a label for every low-quality answer. This shortcut removes useful friction, but it also hides an assumption that should be tested. Replace it with an observable check.
  • Correcting wording while leaving bad evidence untouched. This shortcut removes useful friction, but it also hides an assumption that should be tested. Replace it with an observable check.
  • Assuming one fixed error means the workflow is reliable. This shortcut removes useful friction, but it also hides an assumption that should be tested. Replace it with an observable check.

Read

Build the habit

Before you close the tab, capture a reusable habit for A map of AI mistakes inside Why AI makes mistakes: name the observable check, the evidence you would open, and the stop condition. Rehearse it once on a low-stakes example, then once on a higher-stakes variant. The habit succeeds when you can explain the check without reopening this lesson. Target outcome: Explain why plausible language generation does not guarantee factual accuracy.

Go deeper

Before you start

Why this matters

A project brief contains an invented statistic, an outdated product name, and a correct fact attached to the wrong source.

In the wild

See how this idea shows up as a product and a company — then come back to the lesson. Skills transfer across vendors.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

Local focus for A map of AI mistakes (Why AI makes mistakes): write the smallest test that would falsify a confident claim on this page, name the evidence you would open first, and note who must approve if the cost of being wrong is more than a redo. Keep the note under ten lines so you will actually reuse it.

1. Explain the page’s core distinction without using the word “smart.”
2. Which fact, source, permission, or test would most change your judgment in the opening scenario?
3. Name one low-consequence use where a light check is enough and one high-consequence use where independent review is required.
4. What should a responsible user do when the available evidence cannot support the requested conclusion?

All responses are required.