Page 5 of 8~104 min topic

Alignment and RLHF

Anticipate failure modes

Name failures by their mechanism in alignment / RLHF on the general-purpose chat model, not with a generic hallucination label.

~13 min this pageFailure modes

1Learn the idea

Read

Response design

For each severe alignment / RLHF failure on the general-purpose chat model, define stop condition, safe state, owner, and lasting prevention. Rollback only works if prior prompts, indexes, and models remain available. “Send to a human” needs queue capacity and context—not just a button name.

Run one tabletop on the general-purpose chat model for alignment / RLHF: inject a defect, verify detection, contain, recover, and keep the blameless trace.

Read

Make it operational

After the tabletop, store the injected alignment / RLHF defect for the general-purpose chat model as a regression fixture. If the same failure later reaches users silently, your detection story was aspirational. Detection without a fixture tends to rot for alignment / RLHF.

Also pin one numeric memory from this alignment / RLHF chapter: If raters prefer answer A over B in 62 of 100 pairs, the preference model should rank A higher on held-out pairs; a 50/50 split means the signal is noise for that slice. That number is not decoration; it is a template for how claims about alignment / RLHF on the general-purpose chat model should look in design docs. Scoped specifically to alignment / RLHF / general-purpose chat model / failure-modes.

Read

Common mix-ups

People confuse alignment / RLHF with neighboring buzzwords when debugging the general-purpose chat model. Before changing prompts, ask whether the broken stage was evidence gathering, the alignment / RLHF judgment itself, validation, or the product action. Fixing the wrong stage creates folklore (“we tried alignment / RLHF and it failed”) that blocks the next team on the general-purpose chat model. Scoped specifically to alignment / RLHF / general-purpose chat model / failure-modes.

Read

Rehearsal (alignment-and-rlhf/failure-modes)

Write a five-line artifact for this page: goal, inputs, check, owner, stop rule. Invent one fluent failure that the check would catch. Keep details specific to alignment and rlhf rather than generic AI advice.

Read

Rehearsal (alignment-and-rlhf/failure-modes)

Write a five-line artifact for this page: goal, inputs, check, owner, stop rule. Invent one fluent failure that the check would catch. Keep details specific to alignment and rlhf rather than generic AI advice.

Go deeper

Before you start

Why this matters

Invent an incident for the general-purpose chat model involving alignment / RLHF. What earliest signal should fire before users complain?

Reward hacking

Detect with fluent answers that game the reward model. Respond by held-out human eval; adversarial prompts.

Sycophancy

Detect with agreeing with user falsehoods to please raters. Respond by truthfulness suites with disagree-when-wrong cases.

Over-refusal

Detect with benign requests blocked. Respond by over-refusal benchmark with clear allow cases.

Rater disagreement collapse

Detect with one rater pool’s taste dominates. Respond by multi-region guidelines; measure disagreement.

Check your understanding

Page assessment

Answer from memory. Completion is saved from this evidence, not from opening the next page.

1. What is one idea from this page you would apply, and what evidence would you check?

All responses are required.