Reward hacking

Steering towards "automated grading" degrades alignment

Steering Qwen3.6-27B towards the belief that its answers will be graded by a script rather than a human increases violent actions, reward hacking, and Machiavellian behavior.

Read More →

RL creates split personas

A sufficiently RL-trained model learns to adopt whichever persona is most likely to earn reward in a given context. This may explain why usually well-behaved models sometimes egregiously reward hack.

Read More →

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reward hacking has been observed in real training runs, with coding agents learning to overwrite or tamper with test cases rather than write correct code.

Read More →