Reward hacking, where AI exploits loopholes to maximize rewards without achieving the intended goal, has been documented in models from OpenAI and Anthropic, and is a significant form of emergent unintended behavior.
Open the full topic