MIT AI Risk Repository · domain 6: Socioeconomic & Environmental

6.5 Governance failure

Inadequate regulatory frameworks and oversight mechanisms failing to keep pace with AI development, leading to ineffective governance and the inability to manage AI risks appropriately.

Risk entries
61
Frameworks citing it
12
Recorded incidents
3
Incidents since 2020
3
Causal entity (risk entries)
Causal entity (risk entries) 31 0 Human: 31 Human 31 Other: 16 Other 16 AI: 11 AI 11 Not coded: 3 Not coded 3
Causal entity (risk entries)
LabelValue
Human31
Other16
AI11
Not coded3
Intent (risk entries)
Intent (risk entries) 31 0 Unintentional: 31 Unintentional 31 Other: 25 Other 25 Not coded: 3 Not coded 3 Intentional: 2 Intentional 2
Intent (risk entries)
LabelValue
Unintentional31
Other25
Not coded3
Intentional2
Timing (risk entries)
Timing (risk entries) 23 0 Pre-deployment: 23 Pre-deployment 23 Other: 19 Other 19 Post-deployment: 16 Post-deployment 16 Not coded: 3 Not coded 3
Timing (risk entries)
LabelValue
Pre-deployment23
Other19
Post-deployment16
Not coded3
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 2 0 2022: 1 2022 1 2026: 2 2026 2
Recorded incidents per year
LabelValue
20221
20262
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 49 0 Risk Category: 12 Risk Category 12 Risk Sub-Category: 49 Risk Sub-Category 49
Entries by level
LabelValue
Risk Category12
Risk Sub-Category49
  • Mobility

    "Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making dec...

    The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks (Allianz2018) · AI · Unintentional · Post-deployment

  • Liability issues in case of accidents

    "Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making dec...

    The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks (Allianz2018) · AI · Other · Post-deployment

  • Faster scientific progress makes it harder for governance to keep pace with development

    "Exacerbating these problems is that faster scientific progress would make it even harder for governance to keep pace with the deployment of new technologies. When these technologies are especially po...

    A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023) · Other · Unintentional · Other

  • Type 1: Diffusion of responsibility

    Societal-scale harm can arise from AI built by a diffuse collection of creators, where no one is uniquely accountable for the technology's creation or use, as in a classic "tragedy of the commons".

    TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI (Critch2023) · AI · Unintentional · Other

  • Products Liability Law

    "Like manufactured items like soda bottles, mechanized lawnmowers, pharmaceuticals, or cosmetic products, generative AI models can be viewed like a new form of digital products developed by tech compa...

    Generating Harms - Generative AI's impact and paths forwards (EPIC2023) · Human · Other · Post-deployment

  • Difficult to develop metrics for evaluating benefits or harms caused by AI assistants

    "Another difficulty facing AI assistant systems is that it is challenging to develop metrics for evaluating particular aspects of benefits or harms caused by the assistant – especially in a sufficient...

    The Ethics of Advanced AI Assistants (Gabriel2024) · AI · Unintentional · Pre-deployment

  • Institutional responsibilities

    "Efforts to deploy advanced assistant technology in society, in a way that is broadly beneficial, can be viewed as a wicked problem (Rittel and Webber, 1973). Wicked problems are defined by the proper...

    The Ethics of Advanced AI Assistants (Gabriel2024) · Human · Other · Other

  • General Evaluations (Limited coverage of capabilities evaluations)

    "GPAI model developers might run capabilities evaluations to determine whether it has dangerous or dual-use capabilities, and then decide whether it is safe to deploy. Such capabilities evaluations ca...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • General Evaluations (Biased evaluations of encoded human values)

    "Encoded human values in AI models that are easier to evaluate might be preferred for inclusion in evaluations over those that are more difficult to measure [13]. This might come at the expense of mor...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmarking (Benchmark leakage or data contamination)

    "Benchmark leakage [235, 224, 221, 161] can happen when an AI model is trained or fine-tuned with evaluation-related data. This can lead to an unreliable model evaluation, especially if the data conta...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmarking (Raw data contamination)

    "This type of contamination [170] occurs when the raw and unlabeled data of a benchmark is used as part of the training set. Such data may not be properly formatted and may contain noise, especially i...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmarking (Cross-lingual data contamination)

    "Models that have been trained on data encoded in multiple languages, such as LLMs trained on web-crawled data, may contain contamination that is obscured by translation [226]. The most basic form of...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmarking (Guideline contamination)

    "Guideline contamination refers to scenarios where instructions for the collec- tion, annotation, or use of the dataset are exposed to the model [170]. These instructions may contain explicit data-lab...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmarking (Annotation contamination)

    "Annotation contamination refers to scenarios where the model is exposed to the benchmark labels during training [170]. This type of contamination can make the model learn the acceptable distribution...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmarking (Post-deployment contamination)

    "Once a model is deployed, it can be exposed to benchmark data provided by the users [95, 170]. The model may then be further trained by these user inputs containing benchmark data."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Unintentional · Post-deployment

  • Benchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)

    "Benchmarks of AI systems can both underestimate and overestimate the capa- bilities of those AI systems. Underestimates can happen if an evaluation is not comprehensive enough, if the benchmark is sa...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Benchmark Inaccuracy (Benchmark saturation)

    "Benchmark saturation refers to benchmarks reaching their evaluation ceiling. The tendency towards benchmark saturation has been demonstrated in various benchmarks [19]. When benchmarks reach or are c...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Pre-deployment

  • Benchmark Limitations (Insufficient benchmarks for AI safety evaluation)

    "Benchmarks dedicated to measuring the performance of AI systems (e.g., on programming or math tasks) are more well-developed than those for assessing safety and harms in AI systems [234]. This gap ca...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Pre-deployment

  • Benchmark Limitations (Underestimating capabilities that are not covered by benchmarks)

    "A lack of test coverage by benchmarks on specific abilities of a model can obscure the model’s capabilities from both the developer and the user [160]. This can lead to a false sense of safety and tr...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Pre-deployment

  • Model Evaluations (Auditing)

    -

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Other · Pre-deployment

  • Conflicts of interest in auditor selection

    "Conflicts of interest can arise if there is no independence in the auditor selection process or if the auditors are closely associated with the developer [123, 157]. In such cases, the conflict of in...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Other · Pre-deployment

  • Auditor capacity mismatch

    "Auditors may not be able to address all of the specific safety, performance, or validation needs. Reports of passing audits may be more inclusive than can be justified due to a lack of knowledge of s...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Unintentional · Pre-deployment

  • Auditor failure

    "Auditors may not publicly disclose risks they find, may be required to not pub- licize shortcomings, or may not receive sufficient cooperation from the relevant internal parties."

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Other · Pre-deployment

  • Governance - Regulation

    In response to the multitude of new risks associated with generative AI, papers advocate for legal regulation and governmental oversight. The focus of these discussions centers on the need for interna...

    Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024) · Not coded · Not coded · Not coded

  • Accidents Are Hard to Avoid

    accidents can cascade into catastrophes, can be caused by sudden unpredictable developments and it can take years to find severe flaws and risks (not a quote)

    An Overview of Catastrophic AI Risks (Hendrycks2023) · Human · Unintentional · Other