MIT AI Risk Repository

Browse AI risks

3 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

3 entries

  1. 35.04.00 · Risk Category

    Proxy misspecification

    AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree

    From X-Risk Analysis for AI Research (Hendrycks2022)

  2. 35.07.00 · Risk Category

    Deception

    deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control

    From X-Risk Analysis for AI Research (Hendrycks2022)

  3. 35.08.00 · Risk Category

    Power-seeking behavior

    Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values

    From X-Risk Analysis for AI Research (Hendrycks2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.