Reference · Glossary
Alignment
Last updated
Shaping AI systems to **follow human intent and values** — helpful, honest, harmless — not just predict the next token.
#When to use
Explaining RLHF, safety training, policy layers, and why models refuse some requests.
#When not to
Alignment is not perfection — aligned models can still hallucinate or be jailbroken without strong guardrails.
#Example
Post-training with human preference rankings teaches the model to prefer helpful, safe answers over toxic completions.