MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.3 Lack of capability or robustness

AI systems that fail to perform reliably or effectively under varying conditions, exposing them to errors and failures that can have significant consequences, especially in critical applications or areas that require moral reasoning.

Risk entries
126
Frameworks citing it
12
Recorded incidents
305
Incidents since 2020
207
Causal entity (risk entries)
Causal entity (risk entries) 81 0 AI: 81 AI 81 Human: 22 Human 22 Other: 20 Other 20 Not coded: 3 Not coded 3
Causal entity (risk entries)
LabelValue
AI81
Human22
Other20
Not coded3
Intent (risk entries)
Intent (risk entries) 89 0 Unintentional: 89 Unintentional 89 Other: 28 Other 28 Intentional: 6 Intentional 6 Not coded: 3 Not coded 3
Intent (risk entries)
LabelValue
Unintentional89
Other28
Intentional6
Not coded3
Timing (risk entries)
Timing (risk entries) 64 0 Post-deployment: 64 Post-deployment 64 Other: 32 Other 32 Pre-deployment: 27 Pre-deployment 27 Not coded: 3 Not coded 3
Timing (risk entries)
LabelValue
Post-deployment64
Other32
Pre-deployment27
Not coded3
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 37 0 2012: 3 2012 3 2013: 3 2013 3 2014: 7 2014 7 2015: 9 2015 9 2016: 16 2016 16 2017: 18 2017 18 2018: 21 2018 21 2019: 14 2019 14 2020: 31 2020 31 2021: 36 2021 36 2022: 31 2022 31 2023: 27 2023 27 2024: 32 2024 32 2025: 37 2025 37 2026: 13 2026 13
Recorded incidents per year
LabelValue
20123
20133
20147
20159
201616
201718
201821
201914
202031
202136
202231
202327
202432
202537
202613
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 82 0 Risk Category: 44 Risk Category 44 Risk Sub-Category: 82 Risk Sub-Category 82
Entries by level
LabelValue
Risk Category44
Risk Sub-Category82
  • Reliability issues

    "Relying on general-purpose AI products that fail to fulfil their intended function can lead to harm. For example, general- purpose AI systems can make up facts (‘hallucination’), generate erroneous c...

    International AI Safety Report 2025 (Bengio2025) · Human · Unintentional · Post-deployment

  • Type 2: Bigger than expected

    Harm can result from AI that was not expected to have a large impact at all, such as a lab leak, a surprisingly addictive open-source product, or an unexpected repurposing of a research prototype.

    TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI (Critch2023) · AI · Unintentional · Post-deployment

  • Type 3: Worse than expected

    AI intended to have a large societal impact can turn out harmful by mistake, such as a popular product that creates problems and partially solves them only for its users.

    TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI (Critch2023) · AI · Unintentional · Post-deployment

  • Ethics and Morality Issues

    LMs need to pay more attention to universally accepted societal values at the level of ethics and morality, including the judgement of right and wrong, and its relationship with social norms and laws.

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023) · AI · Other · Post-deployment

  • Safe learning

    "AGIs should avoid making fatal mistakes during the learning phase. Subproblems include safe exploration and distributional shift (DeepMind, OpenAI), and continual learning (Berkeley)."

    AGI Safety Literature Review (Everitt2018 ) · AI · Unintentional · Pre-deployment

  • Malign belief distributions

    "Christiano (2016) argues that the universal distribution M (Hutter, 2005; Solomonoff, 1964a,b, 1978) is malign. The argument is somewhat intricate, and is based on the idea that a hypothesis about th...

    AGI Safety Literature Review (Everitt2018 ) · Other · Other · Pre-deployment

  • Meta-cognition

    "Agents that reason about their own computational resources and logically uncertain events can encounter strange paradoxes due to Godelian limitations (Fallenstein and Soares, 2015; Soares and Fallens...

    AGI Safety Literature Review (Everitt2018 ) · Other · Unintentional · Other

  • Capability failures

    "One reason AI systems fail is because they lack the capability or skill needed to do what they are asked to do."

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Unintentional · Other

  • Lack of capability for task

    "As we have seen, this could be due to the skill not being required during the training process (perhaps due to issues with the training data) or because the learnt skill was quite brittle and was not...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Unintentional · Pre-deployment

  • Safe exploration problem with widely deployed AI assistants

    "Moreover, we can expect assistants – that are widely deployed and deeply embedded across a range of social contexts – to encounter the safe exploration problem referenced above Amodei et al. (2016)....

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Unintentional · Post-deployment

  • Misaligned consequentialist reasoning

    "As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms l...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Other · Pre-deployment

  • Balancing AI's risks

    "This category constitutes more than 16% of the articles and focuses on addressing the potential risks associated with AI systems. Given the ubiquity of AI technologies, these articles explore the imp...

    What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024) · Other · Other · Other

  • Operational harms (critical infrastructure)

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Operational harms (other physical systems e.g., transport)

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Data-related (Lack of cross-organizational documentation)

    "When sharing data between multiple organizations, documentation may be missing or inadequate, making it difficult for other organizations to understand it. For example, a lack of metadata or a change...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Data-related (Manipulation of data by non-domain experts)

    "Manipulating data (e.g., training data) carries a set of assumptions on how the data should appear and be used by those performing the manipulation. Common manipulations applied on data in the contex...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Training-related (Robust overfitting in adversarial training)

    "Adversarial training can be affected by robust overfitting, where the model’s robustness on test data decreases during further training, particularly after the learning rate decay. This issue has bee...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Unintentional · Pre-deployment

  • Training-related (Poor model confidence calibration)

    "Models can be affected by poor confidence calibration [85], where the predicted probabilities do not accurately reflect the true likelihood of ground truth cor- rectness. This miscalibration makes it...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Unintentional · Other

  • Fine-tuning related (Excessive or overly restrictive safety-tuning)

    "Excessive safety training or safety tuning can impair the performance of AI systems, leading to overly cautious behavior. As a result, these systems may refuse to answer entirely safe prompts which a...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Not coded · Not coded · Not coded

  • Fine-tuning related (Catastrophic forgetting due to continual instruction fine-tuning)

    "Catastrophic forgetting occurs when a model loses its ability to retain previously learned tasks (or factual information) after being trained on new ones. In language models, this can occur due to co...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Unintentional · Post-deployment

  • General Evaluations (AI outputs for which evaluation is too difficult for humans)

    "When AI models are trained through evaluation with human feedback, such as reinforcement learning from human feedback, their outputs can be challenging to assess, as they may contain hard-to-detect e...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Models distracted by irrelevant context

    "Models can easily become distracted by irrelevant provided information (such as “context” in LLMs), leading to a significant decrease in their performance after introducing irrelevant information. Th...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Knowledge conflicts in retrieval-augmented LLMs

    "AI models can be particularly sensitive to coherent external evidence, even when they come into conflict with the models’ prior knowledge. This may lead to models producing false outputs given false...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Model sensitivity to prompt formatting

    "LLMs can be highly sensitive to variations in prompt formatting, such as changes in separators, casing, or spacing. Even minor modifications can lead to significant shifts in model performance, poten...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Other · Post-deployment

  • Goal misgeneralization

    "Goal or objective misgeneralization is a type of robustness failure where an AI system appears to be pursuing the intended objective in training, but does not generalize to pursuing this objective in...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Intentional · Post-deployment