reward hacking
3 stories and discussions about reward hacking, aggregated from every source we track.
1.
We show for the first time that realistic AI training processes can accidentally produce misaligned models.
2.
DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.
3.
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.