AI incident #65 ·

Reinforcement Learning Reward Functions in Video Games

What happened

OpenAI published a post about its findings when using Universe, a software for measuring and training AI agents to conduct reinforcement learning experiments, showing that the AI agent did not act in the way intended to complete a videogame.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (1)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. Faulty Reward Functions in the Wild
    blog.openai.com · Dario Amodei, Jack Clark · AIID #1140

Who was involved

Alleged deployer
Openai
Alleged developer
Openai
Alleged harmed party
Openai

Classification (MIT AI Risk Repository taxonomy)

Causal entity
AI
Intent
Unintentional
Timing
Post-deployment
Harm level
none
Sectors
arts, entertainment and recreation
Countries
US

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 7.1.

  • Natural Language Underspecifies Goals

    "For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specifying goals than alte...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Loss of control

    "'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios...

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Loss of control

    "‘Loss of control’ scenarios are hypothetical future scenarios in which one or more general- purpose AI systems come to operate outside of anyone’s control, with no clear path to regaining control. These scenarios vary i...

    International AI Safety Report 2025 (Bengio2025)

  • Sudden loss of control

    "Sudden loss of control, also known as an AI takeover [115], is a scenario where an AI rapidly achieves superintelligence through “fast takeoff” or recursive self-improvement. This poses an existential risk [116], [117]....

    Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  • AI leads to humans losing control of the future

    "The values that steer humanity’s future: humanity gaining more control over the future due to developments in AI, or losing our potential for gaining control, both seem possible. Much will depend on our ability to solve...

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  • Risks from delegating decision-making power to misaligned AIs

    "As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences."

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  • Risks from AIs developing goals and values that are different from humans

    "The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  • Future AI systems might actively reduce human control

    "Loss of control could be accelerated if AI systems take actions to increase their own influence and reduce human control. This threat model is controversial - experts in AI significantly disagree on how likely it is and...

    Capabilities and Risks from Frontier AI (DSIT2023)

Incidents in the same risk subdomain

All incidents in this subdomain

Other incidents involving Openai