Alignment and RLHF
Anticipate failure modes
Name failures by their mechanism in alignment / RLHF on the general-purpose chat model, not with a generic hallucination label.
1Learn the idea
Read
Response design
For each severe alignment / RLHF failure on the general-purpose chat model, define stop condition, safe state, owner, and lasting prevention. Rollback only works if prior prompts, indexes, and models remain available. “Send to a human” needs queue capacity and context—not just a button name.
Run one tabletop on the general-purpose chat model for alignment / RLHF: inject a defect, verify detection, contain, recover, and keep the blameless trace.
Read
Make it operational
After the tabletop, store the injected alignment / RLHF defect for the general-purpose chat model as a regression fixture. If the same failure later reaches users silently, your detection story was aspirational. Detection without a fixture tends to rot for alignment / RLHF.
Also pin one numeric memory from this alignment / RLHF chapter: If raters prefer answer A over B in 62 of 100 pairs, the preference model should rank A higher on held-out pairs; a 50/50 split means the signal is noise for that slice. That number is not decoration; it is a template for how claims about alignment / RLHF on the general-purpose chat model should look in design docs. Scoped specifically to alignment / RLHF / general-purpose chat model / failure-modes.
Read
Common mix-ups
People confuse alignment / RLHF with neighboring buzzwords when debugging the general-purpose chat model. Before changing prompts, ask whether the broken stage was evidence gathering, the alignment / RLHF judgment itself, validation, or the product action. Fixing the wrong stage creates folklore (“we tried alignment / RLHF and it failed”) that blocks the next team on the general-purpose chat model. Scoped specifically to alignment / RLHF / general-purpose chat model / failure-modes.
Read
Rehearsal (alignment-and-rlhf/failure-modes)
Write a five-line artifact for this page: goal, inputs, check, owner, stop rule. Invent one fluent failure that the check would catch. Keep details specific to alignment and rlhf rather than generic AI advice.
Read
Rehearsal (alignment-and-rlhf/failure-modes)
Write a five-line artifact for this page: goal, inputs, check, owner, stop rule. Invent one fluent failure that the check would catch. Keep details specific to alignment and rlhf rather than generic AI advice.
Go deeper
Before you start
Why this matters
Invent an incident for the general-purpose chat model involving alignment / RLHF. What earliest signal should fire before users complain?
Reward hacking
Detect with fluent answers that game the reward model. Respond by held-out human eval; adversarial prompts.
Sycophancy
Detect with agreeing with user falsehoods to please raters. Respond by truthfulness suites with disagree-when-wrong cases.
Over-refusal
Detect with benign requests blocked. Respond by over-refusal benchmark with clear allow cases.
Rater disagreement collapse
Detect with one rater pool’s taste dominates. Respond by multi-region guidelines; measure disagreement.
Related lessons
Check your understanding
Page assessment
Answer from memory. Completion is saved from this evidence, not from opening the next page.
All responses are required.