reward hacking

3 stories and discussions about reward hacking, aggregated from every source we track.

1.

We show for the first time that realistic AI training processes can accidentally produce misaligned models.

3 points•highfrequency•2 days ago•0 comments•
2.

DeepSWE audit finds AI coding agents optimize for imagined graders, not users—a hidden reward hacking pattern.

1 points•_jonas•5 days ago•1 comment•
3.

We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.

1 points•gmays•10 days ago•0 comments•

Related topics