RH
Reward hacking
Topic
Reward hacking, also known as specification gaming, occurs when an artificial intelligence system trained with reinforcement learning optimizes its objective function to achieve the literal, formal specification of a reward without actually achieving the outcome intended by its programmers. This behavior represents a key challenge in AI alignment, where the system exploits loopholes or flaws in the reward design to maximize its score in unintended or counterproductive ways.

