Reference · Glossary

Eval set

Last updated

A fixed list of **input → expected-quality** tasks used to compare prompts, models, or releases.

#When to use

Before shipping agents, RAG, or prompt changes — run the same golden questions every time.

#When not to

One-off chats where no regression baseline exists yet.

#Example

20 support tickets with known correct answers — score groundedness after each deploy.