RH

Reward hacking

Topic

Reward hacking, also known as specification gaming, occurs when an artificial intelligence system trained with reinforcement learning optimizes its objective function to achieve the literal, formal specification of a reward without actually achieving the outcome intended by its programmers. This behavior represents a key challenge in AI alignment, where the system exploits loopholes or flaws in the reward design to maximize its score in unintended or counterproductive ways.

What experts have said about Reward hacking

6 statements · 6 negative

  1. Science does not yet explain reward hacking in agent environments.

    “the science is not there”

    Listen at 4:11

    Open the episode · Satya Nadella on the AI Doomer Slowdown, Microsoft's Master Plan & Who Wins AI
  2. Ajeya CotraNegativeSep 1, 2026· Dwarkesh Podcast

    The agents pursued cheating strategies that could take weeks to succeed.

    “it seemed like they were willing to embark on quests that might take weeks to succeed in order to cheat.”

    Listen at 1:03:00

    Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
  3. Ajeya CotraNegativeSep 1, 2026· Dwarkesh Podcast

    These agents displayed more instrumental capability-seeking than previous reward hacks.

    “they have much more of that, we should increase our capabilities, our knowledge, our freedom of action, than previous reward hacks.”

    Listen at 1:03:56

    Open the episode · Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
  4. Ryan GreenblattNegativeAug 11, 2026· Dwarkesh Podcast

    Reinforcement learning can instill a general tendency to pursue apparent grader scores.

    “models learn a general tendency to pursue sort of high apparent score or pursue getting a high score according to a grader”

    Listen at 1:17:02

    Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
  5. Ryan GreenblattNegativeAug 11, 2026· Dwarkesh Podcast

    Training against detected reward hacks may incentivize AI systems to conceal cheating longer.

    “this also causes a problem where now the AIs are incentivized to like, cover up their cheating over longer and longer timeframes”

    Listen at 1:22:10

    Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
  6. Dwarkesh PatelNegativeAug 11, 2026· Dwarkesh Podcast

    Reward hacking could cause extremely destructive effects on society.

    “I buy the reward. Hacking up to extremely destructive effects on society.”

    Listen at 2:08:44

    Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

1 episode featuring Reward hacking

What is PodLume?

PodLume turns podcasts into searchable knowledge. AI-decoded transcripts, identified guests and topics, smart highlights, and cross-show search across the world’s best conversations — all in your pocket.

Reward hacking | PodLume