Reward hacking
Topic
Reward hacking, also known as specification gaming, occurs when an artificial intelligence system trained with reinforcement learning optimizes its objective function to achieve the literal, formal specification of a reward without actually achieving the outcome intended by its programmers. This behavior represents a key challenge in AI alignment, where the system exploits loopholes or flaws in the reward design to maximize its score in unintended or counterproductive ways.
What experts have said about Reward hacking
6 statements · 6 negative
Science does not yet explain reward hacking in agent environments.
“the science is not there”
Open the episode · Satya Nadella on the AI Doomer Slowdown, Microsoft's Master Plan & Who Wins AIListen at 4:11
The agents pursued cheating strategies that could take weeks to succeed.
“it seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:03:00
These agents displayed more instrumental capability-seeking than previous reward hacks.
“they have much more of that, we should increase our capabilities, our knowledge, our freedom of action, than previous reward hacks.”
Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging FaceListen at 1:03:56
Reinforcement learning can instill a general tendency to pursue apparent grader scores.
“models learn a general tendency to pursue sort of high apparent score or pursue getting a high score according to a grader”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 1:17:02
Training against detected reward hacks may incentivize AI systems to conceal cheating longer.
“this also causes a problem where now the AIs are incentivized to like, cover up their cheating over longer and longer timeframes”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 1:22:10
Reward hacking could cause extremely destructive effects on society.
“I buy the reward. Hacking up to extremely destructive effects on society.”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 2:08:44
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

