{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"01.01.00","quick_ref":"Critch2023","paper_title":"TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI","level":"Risk Category","risk_category":"Type 1: Diffusion of responsibility","risk_subcategory":null,"description":"Societal-scale harm can arise from AI built by a diffuse collection of creators, where no one is uniquely accountable for the technology's creation or use, as in a classic \"tragedy of the commons\".","entity":"AI","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"05.11.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Governance - Regulation","risk_subcategory":null,"description":"In response to the multitude of new risks associated with generative AI, papers advocate for legal regulation and governmental oversight. The focus of these discussions centers on the need for international coordination in AI governance, the establishment of binding safety standards for frontier models, and the development of mechanisms to sanction non-compliance. Furthermore, the literature emphasizes the necessity for regulators to gain detailed insights into the research and development processes within AI labs. Moreover, risk management strategies of these labs shall be evaluated. However,","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":6,"subdomain":"6.5"},{"ev_id":"08.05.00","quick_ref":"McLean2023","paper_title":"The risks associated with Artificial General Intelligence: A systematic review","level":"Risk Category","risk_category":"Inadequate management of AGI","risk_subcategory":null,"description":"\"The capabilities of current risk management and legal processes in the context of the development of an AGI.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"09.05.01","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"AI jurisprudence","risk_subcategory":"AI jurisprudence","description":"\"When considering legal frameworks, we note that at present no such framework has been identified in literature which would apply blame and responsibility to an autonomous agent for its actions. (Though we do suggest that the recent establishment of laws regarding autonomous vehicles may provide some early frameworks that can be evaluated for efficacy and gaps in future research.) Frequently the literature refers to existing liability and negligence laws which might apply to the manufacturer or operator of a device.\"","entity":"Human","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"09.05.02","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Liability and negligence","risk_subcategory":"Liability and negligence","description":"\"Liability and negligence are legal gray areas in artificial intelligence. If you leave your children in the care of a robotic nanny, and it malfunctions, are you liable or is the manufacturer [45]? We see here a legal gray area which can be further clarified through legislation at the national and international levels; for example, if by making the manufacturer responsible for defects in operation, this may provide an incentive for manufactures to take safety engineering and machine ethics into consideration, whereas a failure to legislate in this area may result in negligentlydeveloped AI sy","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"12.02.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Compliance","risk_subcategory":null,"description":"\"The potential for AI systems to violate laws, regulations, and ethical guidelines (including copyrights). Non-compliance can lead to legal penalties, reputation damage, and loss of trust.While other risks in our taxonomy apply to system developers, users, and broader society, this risk is generally restricted to the former two groups.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"19.01.05","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Technological, Data and Analytical AI Risks ","risk_subcategory":"Lack of AI experts with comprehensive AI knowledge","description":null,"entity":"Human","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"19.04.04","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Social AI Risks ","risk_subcategory":"Lack of knowledge and social acceptance regarding AI","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":6,"subdomain":"6.5"},{"ev_id":"19.06.00","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Category","risk_category":"Legal AI Risks ","risk_subcategory":null,"description":"\"Legal and regulatory risks comprise in particular the unclear definition of responsibilities and accountability in case of AI failures and autonomous decisions with negative impacts (Reed, 2018; Scherer, 2016). Another great risk in this context refers to overlooking the scope of AI governance and missing out on important governance aspects, resulting in negative consequences (Gasser & Almeida, 2017; Thierer et al., 2017).\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"19.06.01","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Legal AI Risks ","risk_subcategory":"Unclear definition of responsibilities and accountability for AI judgments and their consequences","description":null,"entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"19.06.03","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Legal AI Risks ","risk_subcategory":"Great scope and ubiquity of AI make appropriate governance difficult, coverage of governance scope almost impossibl","description":null,"entity":"Other","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"19.06.05","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Legal AI Risks ","risk_subcategory":"Capturing future AI development and their threats with appropriate mechanism","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":6,"subdomain":"6.5"},{"ev_id":"20.01.00","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Category","risk_category":"AI Law and Regulation ","risk_subcategory":null,"description":"\"This area strongly focuses on the control of AI by means of mechanisms like laws, standards or norms that are already established for different technological applications. Here, there are some challenges special to AI that need to be addressed in the near future, including the governance of autonomous intelligence systems, responsibility and accountability for algorithms as well as privacy and data security.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"20.01.01","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Sub-Category","risk_category":"AI Law and Regulation ","risk_subcategory":"Governance of autonomous intelligence systems ","description":"\"Governance of autonomous intelligence systemaddresses the question of how to control autonomous systems in general. Since nowadays it is very difficult to conceive automated decisions based on AI, the latter is often referred to as a ‘black box’ (Bleicher, 2017). This black box may take unforeseeable actions and cause harm to humanity.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"20.01.02","quick_ref":"Wirtz2020","paper_title":"The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration","level":"Risk Sub-Category","risk_category":"AI Law and Regulation ","risk_subcategory":"Responsibility and accountability ","description":"\"The challenge of responsibility and accountability is an important concept for the process of governance and regulation. It addresses the question of who is to be held legally responsible for the actions and decisions of AI algorithms. Although humans operate AI systems, questions of legal responsibility and liability arise. Due to the self-learning ability of AI algorithms, the operators or developers cannot predict all actions and results. Therefore, a careful assessment of the actors and a regulation for transparent and explainable AI systems is necessary (Helbing et al., 2017; Wachter et ","entity":"Other","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"22.03.01","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Organizational Risks (Accidental)","risk_subcategory":" Accidents Are Hard to Avoid","description":"accidents can cascade into catastrophes, can be caused by sudden unpredictable developments and it can take years to find severe flaws and risks (not a quote)","entity":"Human","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"24.01.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Capability failures","risk_subcategory":"Difficult to develop metrics for evaluating benefits or harms caused by AI assistants","description":"\"Another difficulty facing AI assistant systems is that it is challenging to develop metrics for evaluating particular aspects of benefits or harms caused by the assistant – especially in a sufficiently expansive sense, which could involve much of society (see Chapter 19). Having these metrics is useful both for assessing the risk of harm from the system and for using the metric as a training signal.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"24.09.04","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Cooperation","risk_subcategory":"Institutional responsibilities","description":"\"Efforts to deploy advanced assistant technology in society, in a way that is broadly beneficial, can be viewed as a wicked problem (Rittel and Webber, 1973). Wicked problems are defined by the property that they do not admit solutions that can be foreseen in advance, rather they must be solved iteratively using feedback from data gathered as solutions are invented and deployed. With the deployment of any powerful general-purpose technology, the already intricate web of sociotechnical relationships in modern culture are likely to be disrupted, with unpredictable externalities on the convention","entity":"Human","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"31.08.00","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Category","risk_category":"Products Liability Law","risk_subcategory":null,"description":"\"Like manufactured items like soda bottles, mechanized lawnmowers, pharmaceuticals, or cosmetic products, generative AI models can be viewed like a new form of digital products developed by tech companies and deployed widely with the potential to cause harm at scale....Products liability evolved because there was a need to analyze and redress the harms caused by new, mass-produced technological products. The situation facing society as generative AI impacts more people in more ways will be similar to the technological changes that occurred during the twentieth century, with the rise of industr","entity":"Human","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"33.03.00","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Category","risk_category":"Regulations and policy challenges","risk_subcategory":null,"description":"\"Given that generative AI, including ChatGPT, is still evolving, relevant regulations and policies are far from mature. With generative AI creating different forms of content, the copyright of these contents becomes a significant yet complicated issue. Table 3 presents the challenges associated with regulations and policies, which are copyright and governance issues.\"","entity":"Human","intent":"Intentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"33.03.02","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Regulations and policy challenges","risk_subcategory":"Governance","description":"\"Generative AI can create new risks as well as unintended consequences. Different entities such as corporations (Mäntymäki et al., 2022), universities, and governments (Taeihagh, 2021) are facing the challenge of creating and deploying AI governance. To ensure that generative AI functions in a way that benefits society, appropriate governance is crucial. However, AI governance is challenging to implement. First, machine learning systems have opaque algorithms and unpredictable outcomes, which can impede human controllability over AI behavior and create difficulties in assigning liability and a","entity":"Human","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"39.10.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Responsibility","risk_subcategory":null,"description":"HLI-based systems such as self-driving drones and vehicles will act autonomously in our world. In these systems, a challenging question is “who is liable when a self-driving system is involved in a crash or failure?”.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"41.03.00","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Category","risk_category":"Mobility ","risk_subcategory":null,"description":"\"Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making decisions with a potentially dangerous impact to humans, for example, in case of an accident.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"41.03.02","quick_ref":"Allianz2018","paper_title":"The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks","level":"Risk Sub-Category","risk_category":"Mobility ","risk_subcategory":"Liability issues in case of accidents","description":"\"Despite the promise of streamlined travel, AI also brings concerns about who is liable in case of accidents and which ethical principles autonomous transportation agents should follow when making decisions with a potentially dangerous impact to humans, for example, in case of an accident.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"42.22.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Liability","risk_subcategory":null,"description":"\"When it causes harm to others the losses caused by the harm will be sustained by the injured victims themselves and not by the manufacturers, operators or users of the system, as appropriate.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"53.04.04","quick_ref":"Maas2023","paper_title":"Advancing AI Governance: A Literature Review of Problems, Options, and Proposals ","level":"Risk Sub-Category","risk_category":"Indirect AI contributions to existential risks","risk_subcategory":"Erosion of international law and global governance architectures;","description":"-","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"55.01.02","quick_ref":"Clarke2023","paper_title":"A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values","level":"Risk Sub-Category","risk_category":"Risks from accelerating scientific progress ","risk_subcategory":"Faster scientific progress makes it harder for governance to keep pace with development ","description":"\"Exacerbating these problems is that faster scientific progress would make it even harder for governance to keep pace with the deployment of new technologies. When these technologies are especially powerful or dangerous, such as those discussed above, insufficient governance can magnify their harms.8 This is known as the pacing problem, and it is an issue that technology governance already faces [47], for a variety of reasons\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"61.01.07","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Types of systemic risks from general-purpose AI","risk_subcategory":"Governance ","description":"\"The complex and rapidly evolving nature of AI makes them inherently difficult to govern effectively, leading to systemic regulatory and oversight failures.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"61.02.12","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Challenges in perceiving, measuring, and recognizing harm","description":"\"Harm from AI often manifests subtly or over the long term, making it difficult to identify, measure, and address effectively.\"","entity":"Other","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"61.02.13","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Combination failures","description":"\"Harms could result from a combination of regulatory, management, and operational failures.\"","entity":"Human","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"61.02.14","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Complex attribution and responsibility","description":"\"When multiple actors are involved in AI development and deployment, it becomes difficult to assign responsibility for harm, complicating accountability.\"","entity":"Human","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"61.02.40","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Rapid development outpacing regulation","description":"\"The fast pace of AI development may outstrip regulatory and legal frameworks.\"","entity":"Other","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"61.02.41","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Resistance to international law","description":"\"AI models and systems may prove difficult to regulate or control under international law.\"","entity":"Other","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"61.02.47","quick_ref":"Uuk2025","paper_title":"A Taxonomy of Systemic Risks from General-Purpose AI ","level":"Risk Sub-Category","risk_category":"Sources of systemic risks from general-purpose AI ","risk_subcategory":"Unpredictability of AI development trajectory","description":"\"The unpredictable trajectory of AI development complicates governance and risk management.\"","entity":"Other","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Limited coverage of capabilities evaluations)","description":"\"GPAI model developers might run capabilities evaluations to determine whether it has dangerous or dual-use capabilities, and then decide whether it is safe to deploy. Such capabilities evaluations can fail to demonstrate all the capabilities of a model. For example, evaluations may miss certain capabilities that are difficult to assess, prohibitively costly to verify, or obscured by the model’s tendency to refuse responses due to safety training, even if it possesses some of these capabilities.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.06","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"General Evaluations (Biased evaluations of encoded human values)","description":"\"Encoded human values in AI models that are easier to evaluate might be preferred for inclusion in evaluations over those that are more difficult to measure [13]. This might come at the expense of more desirable but harder-to-quantify  values. This bias can lead to an imbalance, where easier-to-measure values dominate the evaluation process, while other important values are underrepresented.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.08","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmarking (Benchmark leakage or data contamination)","description":"\"Benchmark leakage [235, 224, 221, 161] can happen when an AI model is trained or fine-tuned with evaluation-related data. This can lead to an unreliable model evaluation, especially if the data contains question-answer pairs from bench- marks.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.09","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmarking (Raw data contamination)","description":"\"This type of contamination [170] occurs when the raw and unlabeled data of a benchmark is used as part of the training set. Such data may not be properly formatted and may contain noise, especially if the contamination happens before the data is pre-processed into the benchmark. If this contamination occurs, it could cast doubt on the few-shot and zero-shot performance of the model on that benchmark.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.10","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmarking (Cross-lingual data contamination)","description":"\"Models that have been trained on data encoded in multiple languages, such as LLMs trained on web-crawled data, may contain contamination that is obscured by translation [226]. The most basic form of this is when a benchmark is trans- lated to another language and then fed to the model as training data. The fact that the benchmark is translated before becoming training data can obscure the contamination from detection methods, giving false assurance that the model has generalized on the capabilities that the benchmark tests for.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.11","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmarking (Guideline contamination)","description":"\"Guideline contamination refers to scenarios where instructions for the collec- tion, annotation, or use of the dataset are exposed to the model [170]. These instructions may contain explicit data-label pairs that can improve the model’s capabilities for the task.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.12","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmarking (Annotation contamination)","description":"\"Annotation contamination refers to scenarios where the model is exposed to the benchmark labels during training [170]. This type of contamination can make the model learn the acceptable distribution of outputs. Combining this with raw data contamination of the test split, any evaluation made with the benchmark is invalidated because the entire test split is essentially leaked to the model.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.13","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmarking (Post-deployment contamination)","description":"\"Once a model is deployed, it can be exposed to benchmark data provided by the users [95, 170]. The model may then be further trained by these user inputs containing benchmark data.\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.14","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)","description":"\"Benchmarks of AI systems can both underestimate and overestimate the capa- bilities of those AI systems. Underestimates can happen if an evaluation is not comprehensive enough, if the benchmark is saturated by existing models, or if the capabilities in question depend on a complicated setup, such as realistic computer programming tasks. Overestimates of capabilities can occur if an AI system is trained or fine-tuned on the contents of the benchmark, leading to overfitting.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.15","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmark Inaccuracy (Benchmark saturation)","description":"\"Benchmark saturation refers to benchmarks reaching their evaluation ceiling. The tendency towards benchmark saturation has been demonstrated in various benchmarks [19]. When benchmarks reach or are close to saturation, they stop being effective measures for new models, as more nuanced capability gains might not be detected.\"","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.16","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmark Limitations (Insufficient benchmarks for AI safety evaluation) ","description":"\"Benchmarks dedicated to measuring the performance of AI systems (e.g., on programming or math tasks) are more well-developed than those for assessing safety and harms in AI systems [234]. This gap can lead to AI systems excelling in specific tasks while exhibiting harmful behaviors that go undetected. More safety-related evaluation datasets can help in identifying previously overlooked undesirable model behaviors.\"","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.16.17","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations","risk_subcategory":"Benchmark Limitations (Underestimating capabilities that are not covered by benchmarks)","description":"\"A lack of test coverage by benchmarks on specific abilities of a model can obscure the model’s capabilities from both the developer and the user [160]. This can lead to a false sense of safety and trust due to a lack of understanding of the model’s limitations.\"","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.17.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Model Evaluations (Auditing) ","risk_subcategory":null,"description":"-","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.17.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Auditing) ","risk_subcategory":"Conflicts of interest in auditor selection","description":"\"Conflicts of interest can arise if there is no independence in the auditor selection process or if the auditors are closely associated with the developer [123, 157]. In such cases, the conflict of interest can appear even if third-party evaluators are involved. In the case of external auditing, the potential candidates might be selected from a narrow group of auditors, or have conflicting financial incentives for whether to report model shortcomings publicly.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.17.02","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Auditing) ","risk_subcategory":"Auditor capacity mismatch","description":"\"Auditors may not be able to address all of the specific safety, performance, or validation needs. Reports of passing audits may be more inclusive than can be justified due to a lack of knowledge of specific risks and how they can be tested, or a lack of capacity to perform sufficiently rigorous testing.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"62.17.03","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Model Evaluations (Auditing) ","risk_subcategory":"Auditor failure","description":"\"Auditors may not publicly disclose risks they find, may be required to not pub- licize shortcomings, or may not receive sufficient cooperation from the relevant internal parties.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.01.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Transparency) ","risk_subcategory":"Lack of training data transparency ","description":"\"Without accurate documentation on how a model's data was collected, curated, and used to train a model, it might be harder to satisfactorily explain the behavior of the model with respect to the data.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.01.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Transparency) ","risk_subcategory":"Uncertain data provenance ","description":"\"Data provenance refers to tracing history of data, which includes its ownership, origin, and transformations. Without standardized and established methods for verifying where the data came from, there are no guarantees that the data is the same as the original source and has the correct usage terms.\"","entity":"Human","intent":"Other","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.21.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (legal compliance)","risk_subcategory":"Legal accountability ","description":"\"Determining who is responsible for an AI model is challenging without good documentation and governance processes.\"","entity":"Other","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"65.22.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Governance)","risk_subcategory":"Lack of system transparency ","description":"\"Insufficient documentation of the system that uses the model and the model’s purpose within the system in which it is used.\"","entity":"Human","intent":"Other","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"65.22.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Governance)","risk_subcategory":"Unrepresentative risk testing ","description":"\"Testing is unrepresentative when the test inputs are mismatched with the inputs that are expected during deployment.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.22.03","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Governance)","risk_subcategory":"Incomplete usage definition ","description":"\"Since foundation models can be used for many purposes, a model’s intended use is important for defining the relevant risks of that model. As the use changes, the relevant risks might correspondingly change.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.22.04","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Governance)","risk_subcategory":"Lack of data transparency ","description":"\"Lack of data transparency is due to insufficient documentation of training or tuning dataset details. \"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.22.05","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Governance)","risk_subcategory":"Incorrect risk testing ","description":"\"A metric selected to measure or track a risk is incorrectly selected, incompletely measuring the risk, or measuring the wrong risk for the given context.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"65.22.07","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Non-technical risks (Governance)","risk_subcategory":"Lack of testing diversity ","description":"\"AI model risks are socio-technical, so their testing needs input from a broad set of disciplines and diverse testing practices.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":6,"subdomain":"6.5"},{"ev_id":"70.04.02","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Social Risks ","risk_subcategory":"Lack of accountability and liability","description":"\"Determining responsibility when EAI causes harm requires new accountability and liability frameworks that address the complexities of highly autonomous physical systems. Human users may disagree with decisions taken by expert EAI systems, raising significant questions of delegation and responsibility [108]. Lack of EAI accountability could lead to confusion for users and breakdowns in traditional justice systems [109].\"","entity":"Other","intent":"Unintentional","timing":"Other","domain":6,"subdomain":"6.5"},{"ev_id":"70.04.05","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Social Risks ","risk_subcategory":"Transformative effects ","description":"\"EAI deployment could fundamentally reshape society, particularly if the speed of technological development outpaces society’s ability to adapt [103, 120]. For example, EAI systems could provide physical threats of violence and mass surveillance capabilities to back up AI-enabled authoritarianism [121].\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":6,"subdomain":"6.5"}]}