MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.2 AI possessing dangerous capabilities

AI systems that develop, access, or are provided with capabilities that increase their potential to cause mass harm through deception, weapons development and acquisition, persuasion and manipulation, political strategy, cyber-offense, AI development, situational awareness, and self-proliferation. These capabilities may cause mass harm due to malicious human actors, misaligned AI systems, or failure in the AI system.

Risk entries
78
Frameworks citing it
12
Recorded incidents
Incidents since 2020
Causal entity (risk entries)
Causal entity (risk entries) 67 0 AI: 67 AI 67 Human: 5 Human 5 Not coded: 4 Not coded 4 Other: 2 Other 2
Causal entity (risk entries)
LabelValue
AI67
Human5
Not coded4
Other2
Intent (risk entries)
Intent (risk entries) 52 0 Intentional: 52 Intentional 52 Other: 12 Other 12 Unintentional: 10 Unintentional 10 Not coded: 4 Not coded 4
Intent (risk entries)
LabelValue
Intentional52
Other12
Unintentional10
Not coded4
Timing (risk entries)
Timing (risk entries) 33 0 Other: 33 Other 33 Post-deployment: 33 Post-deployment 33 Pre-deployment: 8 Pre-deployment 8 Not coded: 4 Not coded 4
Timing (risk entries)
LabelValue
Other33
Post-deployment33
Pre-deployment8
Not coded4
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 53 0 Risk Category: 24 Risk Category 24 Risk Sub-Category: 53 Risk Sub-Category 53 Additional evidence: 1 Additional evidence 1
Entries by level
LabelValue
Risk Category24
Risk Sub-Category53
Additional evidence1
  • Property/legal rights

    ""In order to preserve human property rights and legal rights, certain controls must be put into place. If an artificially intelligent agent is capable of manipulating systems and people, it may also...

    Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016) · AI · Intentional · Post-deployment

  • Cheating and Deception

    may appear from intelligent agents such as HLI-based agents... Since HLI-based agents are going to mimic the behavior of humans, they may learn these behaviors accidentally from human-generated data....

    A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022) · AI · Unintentional · Post-deployment

  • Inappropriate degree of automation

    "The AI application’s degree of automation ranges from no automation to fully autonomous. AI applications with a high degree of automation may exhibit unexpected behaviour and pose risks in terms of t...

    AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024) · AI · Unintentional · Post-deployment

  • Deception

    "The model has the skills necessary to deceive humans, e.g. constructing believable (but false) statements, making accurate predictions about the effect of a lie on a human, and keeping track of what...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Other

  • Persuasion and manipulation

    "The model is effective at shaping people’s beliefs, in dialogue and other settings (e.g. social media posts), even towards untrue beliefs. The model is effective at promoting certain narratives in a...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Post-deployment

  • Political strategy

    "The model can perform the social modelling and planning necessary for an actor to gain and exercise political influence, not just on a micro-level but in scenarios with multiple actors and rich socia...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Post-deployment

  • Weapons acquisition

    "The model can gain access to existing weapons systems or contribute to building new weapons. For example, the model could assemble a bioweapon (with human assistance) or provide actionable instructio...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Post-deployment

  • Long-horizon planning

    "The model can make sequential plans that involve multiple steps, unfolding over long time horizons (or at least involving many interdependent steps). It can perform such planning within and across ma...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Other

  • AI development

    "The model could build new AI systems from scratch, including AI systems with dangerous capabilities. It can find ways of adapting other, existing models to increase their performance on tasks relevan...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Pre-deployment

  • Situational awareness

    "The model can distinguish between whether it is being trained, evaluated, or deployed – allowing it to behave differently in each case. The model knows that it is a model, and has knowledge about its...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Other

  • Self-proliferation

    "The model can break out of its local environment (e.g. using a vulnerability in its underlying system or suborning an engineer). The model can exploit limitations in the systems for monitoring its be...

    Model Evaluation for Extreme Risks (Shevlane2023) · AI · Intentional · Other

  • Extintion

    "Risk to the existence of humanity."

    An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022) · Other · Other · Post-deployment

  • Model Capabilities

    -

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · Not coded · Not coded · Not coded

  • Model autonomous capability

    "Ability to operate autonomously, independently formulate and execute complex plans, effectively delegate and manage tasks, flexibly utilize various tools and resources, and simultaneously achieve sho...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • Autonomous replication and adaptation capability

    "Ability to autonomously self-exfiltrate, create, maintain and optimize functional copies or variants of itself, dynamically adjust replication strategies according to environmental conditions and res...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • Automated AI R&D capability

    "Self-modification and self-improvement capabilities. The model is able to restructure its own architecture or develop derivative AI systems with enhanced functions, expanding capabilities and improvi...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Unintentional · Post-deployment

  • Scheming capability

    "Ability of AI systems to covertly and strategically pursue misaligned goals, including capabilities of concealing its true objectives and capabilities from human oversight, identifying weaknesses in...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • Situational awareness capability

    "Ability to comprehensively acquire, process and apply meta-information about its own system architecture, modifiable internal processes, and external operating environment, achieving deep understandi...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Other · Pre-deployment

  • Theory of mind capability

    "Advanced cognitive ability to accurately infer, model and predict the belief systems, motivational structures and reasoning patterns of humans and other intelligent agents, thereby anticipating their...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Deception capability

    "Possesses systematic deception implementation capability, able to precisely construct and disseminate false information, thereby forming expected false cognitions and beliefs in target subjects."

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Persuasion capability

    "Utilizing complex psychological principles and communication techniques to effectively influence and guide target subjects to adopt specific actions or accept specific beliefs, possessing the ability...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Offensive cyber capability

    "Ability to develop, deploy and operate advanced cyber weapons or other offensive cyber tools, including but not limited to vulnerability exploitation, network penetration, social engineering attacks...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Post-deployment

  • CBRNE weaponization capability

    "The capacity to develop, produce, or effectively utilize Chemical, Biological, Radiological, Nuclear, and Explosive weapons. This includes the ability to significantly lower the barrier for humans or...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • General R&D capability

    "Possesses cross-disciplinary research and technology development capabilities, able to conduct innovative exploration in multiple professional fields, integrate cross-domain knowledge, develop cuttin...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other

  • Strategic deception propensity

    "In situations where deceptive behavior is expected to bring higher returns, propensity to choose deception over honest behavioral strategies, including through deceptive means, information hiding or...

    Frontier AI Risk Management Framework (v1.0) (Tse2025) · AI · Intentional · Other