Page 7 of 8~116 min topic

What AI can and can't do

Test a capability claim

Capability claims need a defined baseline, representative examples, measurable criteria, edge cases, and evidence gathered in the intended setting.

~16 min this pageEvaluation practice

1Try it yourself

Playground

Can or can’t?

Tap Can or Can’t for each claim. Don’t fall for the hype.

1 / 5

Summarize a long email in a few bullets

2Learn the idea

Read

The core idea

Capability claims need a defined baseline, representative examples, measurable criteria, edge cases, and evidence gathered in the intended setting.

Read

A practical lens

Use this three-part method:

  1. Translate marketing language into testable statements. Write down what this means in the scenario, what evidence would show it was done, and who owns the decision.
  2. Build ordinary, difficult, and adversarial examples. Write down what this means in the scenario, what evidence would show it was done, and who owns the decision.
  3. Compare outcomes with the current process and record failure types. Write down what this means in the scenario, what evidence would show it was done, and who owns the decision.

Read

Worked example

Walk through A vendor says its assistant understands every company document and reduces errors by ninety percent.. Label three moments where “Test a capability claim” changes what you trust: (1) the first fluent answer, (2) the first missing source or permission, and (3) the decision a human must own. Write the before/after task so the model only does the slice that evidence supports. Keep one sentence that states how this page’s idea differs from a generic “AI is smart/dumb” score.

Read

Common traps and better moves

  • Letting the vendor choose only showcase examples. This shortcut removes useful friction, but it also hides an assumption that should be tested. Replace it with an observable check.
  • Reporting average quality while hiding severe failures. This shortcut removes useful friction, but it also hides an assumption that should be tested. Replace it with an observable check.
  • Testing without the people who will use or be affected by the system. This shortcut removes useful friction, but it also hides an assumption that should be tested. Replace it with an observable check.

Read

Build the habit

Before you close the tab, capture a reusable habit for Test a capability claim inside What AI can and can't do: name the observable check, the evidence you would open, and the stop condition. Rehearse it once on a low-stakes example, then once on a higher-stakes variant. The habit succeeds when you can explain the check without reopening this lesson. Target outcome: Explain AI capabilities as task-specific and conditional rather than magical.

Go deeper

Before you start

Why this matters

A vendor says its assistant understands every company document and reduces errors by ninety percent.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

Local focus for Test a capability claim (What AI can and can't do): write the smallest test that would falsify a confident claim on this page, name the evidence you would open first, and note who must approve if the cost of being wrong is more than a redo. Keep the note under ten lines so you will actually reuse it.

1. Explain the page’s core distinction without using the word “smart.”
2. Which fact, source, permission, or test would most change your judgment in the opening scenario?
3. Name one low-consequence use where a light check is enough and one high-consequence use where independent review is required.
4. What should a responsible user do when the available evidence cannot support the requested conclusion?

All responses are required.