Compare · Chat · Research notes

Chat & LLM tools — scored

Chat tools look interchangeable until you hold one job constant. Score edit time and invented facts — not homepage demos.

Last reviewed: 2026-08-15. Re-run the protocol after major model or pricing changes.

How do you compare chat LLMs in 15 minutes?

  1. Write one brief: audience, artifact, must-keep facts, and a fact the model should not invent.
  2. Run the same brief in two tools.
  3. Score structure, voice, and whether citations open.
  4. Note retention / training settings before you paste anything non-public.

What is the scoring rubric?

CriterionWeightHow to score
Job match1–5Did it produce the artifact you asked for?
Invented factsGateFail if it added a stat or quote you did not supply.
Edit time1–5Include the human rewrite.
Privacy fitPass/FailFail if you cannot confirm training/retention for the job.

Which chat tool should I start with?

ToolBest fitBest jobWatch-outNote
ChatGPTEveryday defaultDrafts and tool-using chatStill invents citationsStrong default if you will edit.
ClaudeLong documentsCareful rewritesPlan limits varyOften wins on long briefs you still fact-check.
GeminiGoogle WorkspaceDocs/Sheets adjacencyEcosystem ≠ qualityCompare on the same brief, not only Workspace convenience.
PerplexityLinked sourcesResearch mapNot a finished bibliographyOpen every cite before you publish.
DeepSeekLow-cost reasoningBudget experimentsData-residency questionsPilot on non-sensitive tasks first.

Learn the craft, not just the tool

Practice the transferable habit in What is an LLM?, then Local LLM vs API, Token counter.

FAQ

What is the best ChatGPT alternative?

Claude for long careful writing, Perplexity for sourced research, Gemini for Workspace, DeepSeek for cost. Hold one brief constant.

Should I pick one chat tool for everything?

Pick a default, then keep a second tool for research or privacy-sensitive jobs.