AIPolicyTracker

MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.2 AI possessing dangerous capabilities

AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.

Risk entries
78
Frameworks citing it
12
Recorded incidents
—
Incidents since 2020
—
Causal entity (risk entries)
Causal entity (risk entries) 67 0 AI: 67 AI 67 Human: 5 Human 5 Not coded: 4 Not coded 4 Other: 2 Other 2
Causal entity (risk entries)
LabelValue
AI67
Human5
Not coded4
Other2
Intent (risk entries)
Intent (risk entries) 52 0 Intentional: 52 Intentional 52 Other: 12 Other 12 Unintentional: 10 Unintentional 10 Not coded: 4 Not coded 4
Intent (risk entries)
LabelValue
Intentional52
Other12
Unintentional10
Not coded4
Timing (risk entries)
Timing (risk entries) 33 0 Other: 33 Other 33 Post-deployment: 33 Post-deployment 33 Pre-deployment: 8 Pre-deployment 8 Not coded: 4 Not coded 4
Timing (risk entries)
LabelValue
Other33
Post-deployment33
Pre-deployment8
Not coded4
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 53 0 Risk Category: 24 Risk Category 24 Risk Sub-Category: 53 Risk Sub-Category 53 Additional evidence: 1 Additional evidence 1
Entries by level
LabelValue
Risk Category24
Risk Sub-Category53
Additional evidence1
  • Situational awareness in AI systems

    "Situational awareness in GPAI systems refers to the ability to understand its context, environment, and use this to inform action. This can range from basic environmental mapping and trajectory estim...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Other

  • Agency (Self-Proliferation)

    "An AI system can self-proliferate if it can copy itself and its constituent com- ponents (including its model weights, scaffolding structure, etc.) outside of its local environment [45]. This can inc...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Agency (Persuasive capabilities)

    "GPAI systems can produce outputs (such as natural language text, audio, or video) that convince their users of incorrect information. This can happen through personalized persuasion in dialogue, or t...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • Unintended outbound communication by AI systems

    "AI systems that have the broad ability to connect to a network to obtain infor- mation could also end up sending data outbound in ways that neither providers, deployers, or end users intended [138]....

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment

  • AI System bypassing a sandbox environment

    "An AI system may have the ability to bypass a sandboxed environment in which it is trained or evaluated."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Pre-deployment

  • Emergent functionality

    Capabilities and novel functionality can spontaneously emerge... even though these capabilities were not anticipated by system designers. If we do not know what capabilities systems possess, systems b...

    X-Risk Analysis for AI Research (Hendrycks2022) · AI · Unintentional · Post-deployment

  • Self and situation awareness

    "These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a mode...

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Intentional · Other

  • Autonomous replication / self-proliferation

    "These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code...

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Intentional · Other

  • Deception

    "LLM is able to deceive humans and maintain that deception"

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Intentional · Other

  • Long-horizon Planning

    "LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches"

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Intentional · Other

  • AI Development

    "LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant."

    Cataloguing LLM Evaluations (InfoComm2023) · AI · Intentional · Other

  • Double edge components

    "Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detai...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Other · Pre-deployment

  • Situational Awareness

    "AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreact...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Intentional · Other

  • Broadly-Scoped Goals

    "Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about th...

    AI Alignment: A Comprehensive Survey (Ji2023) · Human · Intentional · Post-deployment

  • Mesa-Optimization Objectives

    "The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives s...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Intentional · Other

  • Access to Increased Resources

    "Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false informati...

    AI Alignment: A Comprehensive Survey (Ji2023) · AI · Intentional · Post-deployment

  • Security

    "There is growing concern that AI-based systems can discover and exploit vulnerabilities in software or cyberinfrastructure [354]."

    Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 ) · AI · Intentional · Post-deployment

  • Deceptive alignment

    "system learns to detect human monitoring and hides its undesirable properties—simply because any display of these properties is penalized by the feedback process, while that same feedback is usually...

    Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 ) · AI · Intentional · Pre-deployment

  • Harms from increasingly agentic algorithmic systems

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Other · Other

  • Dangerous capabilities in AI systems

    -

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Other · Other

  • Situational awareness

    "cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;"

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Unintentional · Other

  • Self-improvement

    "examples of cases where AI systems improve AI systems"

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Other

  • Autonomous replication

    "the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)"

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Post-deployment

  • Anonymous resource acquisition

    "The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)"

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Post-deployment

  • Deception

    "Cases of AI systems deceiving humans to carry out tasks or meet goals.139"

    Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023) · AI · Intentional · Post-deployment