MIT AI Risk Repository · domain 6: Socioeconomic & Environmental
6.5 Governance failure
Inadequate regulatory frameworks and oversight mechanisms failing to keep pace with AI development, leading to ineffective governance and the inability to manage AI risks appropriately.
- 61
- 12
- 3
- 3
| Label | Value |
|---|---|
| Human | 31 |
| Other | 16 |
| AI | 11 |
| Not coded | 3 |
| Label | Value |
|---|---|
| Unintentional | 31 |
| Other | 25 |
| Not coded | 3 |
| Intentional | 2 |
| Label | Value |
|---|---|
| Pre-deployment | 23 |
| Other | 19 |
| Post-deployment | 16 |
| Not coded | 3 |
| Label | Value |
|---|---|
| 2022 | 1 |
| 2026 | 2 |
| Label | Value |
|---|---|
| Risk Category | 12 |
| Risk Sub-Category | 49 |
Risk entries
Browse and export all- Mobility
"Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making dec...
- Liability issues in case of accidents
"Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making dec...
- Faster scientific progress makes it harder for governance to keep pace with development
"Exacerbating these problems is that faster scientific progress would make it even harder for governance to keep pace with the deployment of new technologies. When these technologies are especially po...
- Type 1: Diffusion of responsibility
Societal-scale harm can arise from AI built by a diffuse collection of creators, where no one is uniquely accountable for the technology's creation or use, as in a classic "tragedy of the commons".
- Products Liability Law
"Like manufactured items like soda bottles, mechanized lawnmowers, pharmaceuticals, or cosmetic products, generative AI models can be viewed like a new form of digital products developed by tech compa...
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
"Another difficulty facing AI assistant systems is that it is challenging to develop metrics for evaluating particular aspects of benefits or harms caused by the assistant – especially in a sufficient...
- Institutional responsibilities
"Efforts to deploy advanced assistant technology in society, in a way that is broadly beneficial, can be viewed as a wicked problem (Rittel and Webber, 1973). Wicked problems are defined by the proper...
- General Evaluations (Limited coverage of capabilities evaluations)
"GPAI model developers might run capabilities evaluations to determine whether it has dangerous or dual-use capabilities, and then decide whether it is safe to deploy. Such capabilities evaluations ca...
- General Evaluations (Biased evaluations of encoded human values)
"Encoded human values in AI models that are easier to evaluate might be preferred for inclusion in evaluations over those that are more difficult to measure [13]. This might come at the expense of mor...
- Benchmarking (Benchmark leakage or data contamination)
"Benchmark leakage [235, 224, 221, 161] can happen when an AI model is trained or fine-tuned with evaluation-related data. This can lead to an unreliable model evaluation, especially if the data conta...
- Benchmarking (Raw data contamination)
"This type of contamination [170] occurs when the raw and unlabeled data of a benchmark is used as part of the training set. Such data may not be properly formatted and may contain noise, especially i...
- Benchmarking (Cross-lingual data contamination)
"Models that have been trained on data encoded in multiple languages, such as LLMs trained on web-crawled data, may contain contamination that is obscured by translation [226]. The most basic form of...
- Benchmarking (Guideline contamination)
"Guideline contamination refers to scenarios where instructions for the collec- tion, annotation, or use of the dataset are exposed to the model [170]. These instructions may contain explicit data-lab...
- Benchmarking (Annotation contamination)
"Annotation contamination refers to scenarios where the model is exposed to the benchmark labels during training [170]. This type of contamination can make the model learn the acceptable distribution...
- Benchmarking (Post-deployment contamination)
"Once a model is deployed, it can be exposed to benchmark data provided by the users [95, 170]. The model may then be further trained by these user inputs containing benchmark data."
- Benchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)
"Benchmarks of AI systems can both underestimate and overestimate the capa- bilities of those AI systems. Underestimates can happen if an evaluation is not comprehensive enough, if the benchmark is sa...
- Benchmark Inaccuracy (Benchmark saturation)
"Benchmark saturation refers to benchmarks reaching their evaluation ceiling. The tendency towards benchmark saturation has been demonstrated in various benchmarks [19]. When benchmarks reach or are c...
- Benchmark Limitations (Insufficient benchmarks for AI safety evaluation)
"Benchmarks dedicated to measuring the performance of AI systems (e.g., on programming or math tasks) are more well-developed than those for assessing safety and harms in AI systems [234]. This gap ca...
- Benchmark Limitations (Underestimating capabilities that are not covered by benchmarks)
"A lack of test coverage by benchmarks on specific abilities of a model can obscure the model’s capabilities from both the developer and the user [160]. This can lead to a false sense of safety and tr...
- Model Evaluations (Auditing)
-
- Conflicts of interest in auditor selection
"Conflicts of interest can arise if there is no independence in the auditor selection process or if the auditors are closely associated with the developer [123, 157]. In such cases, the conflict of in...
- Auditor capacity mismatch
"Auditors may not be able to address all of the specific safety, performance, or validation needs. Reports of passing audits may be more inclusive than can be justified due to a lack of knowledge of s...
- Auditor failure
"Auditors may not publicly disclose risks they find, may be required to not pub- licize shortcomings, or may not receive sufficient cooperation from the relevant internal parties."
- Governance - Regulation
In response to the multitude of new risks associated with generative AI, papers advocate for legal regulation and governmental oversight. The focus of these discussions centers on the need for interna...
- Accidents Are Hard to Avoid
accidents can cascade into catastrophes, can be caused by sudden unpredictable developments and it can take years to find severe flaws and risks (not a quote)