MIT AI Risk Repository

Browse AI risks

594 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

594 entries · page 1 of 12

  1. "Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  2. 10.02.00 · Risk Category

    Risk of Injury

    "Poorly designed intelligent systems can cause moral, psychological, and physical harm. For example, the use of predictive policing tools may cause more people to be arrested or physically harmed by the police."

    From Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study (Paes2023)

  3. 18.02.02 · Risk Sub-Category

    Misinformation Harms

    Erosion of trust in public information

    "Eroding trust in public information and knowledge"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  4. 19.01.03 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Lack of data, poor data quality, and biases in training data

  5. 37.01.01 · Risk Sub-Category

    Design of AI

    Algorithm and data

    "More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and algorithm opacity."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  6. 45.01.02 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of bias and discrimination)

    "During the algorithm design and training process, personal biases may be introduced, either intentionally or unintentionally. Additionally, poor-quality datasets can lead to biased or discriminatory outcomes in the algorithm's design and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region."

    From AI Safety Governance Framework (TC2602024)

  7. 47.02.10 · Risk Sub-Category

    Ethical and social risks

    Bias and discrimination (value lock and outcome homogenization)

    "Because models are not necessarily retrained to reflect evolving societal views, language models risk “value lock- ins,” which “reifies older, less inclusive understandings.”370 Therefore, the continued use of outdated models may limit the presentation or exploration of alternative perspectives. Moreover, the deployment of identical foundation models by various downstream deployers poses a risk of “outcome homogenization,” creating a potential for homogeneity of bias across broad swathes of society. Identical and widely deployed models with prejudicial training datasets could further entrench

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  8. 50.01.04 · Risk Sub-Category

    System and Operational Risks

    Operational misuses (Automated decision-making)

  9. "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  10. 60.02.02 · Risk Sub-Category

    Risks from malfunctions

    Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of human identity. This can lead to discriminatory outcomes including unequal resource allocation, reinforcement of stereotypes, and systematic neglect of certain groups or viewpoints."

    From International AI Safety Report 2025 (Bengio2025)

  11. 65.04.01 · Risk Sub-Category

    Training Data Risks (Fairness)

    Data bias

    "Historical and societal biases that are present in the data are used to train and fine-tune the model."

    From AI Risk Atlas (IBM2025)

  12. "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  13. Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  14. 30.06.03 · Risk Sub-Category

    Social Norm

    Cultural Insensitivity

    it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  15. 43.02.15 · Risk Sub-Category

    Undesirable Use Cases

    Adult content

    "These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)"

    From Cataloguing LLM Evaluations (InfoComm2023)

  16. 45.01.08 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of improper content and poisoning in training data)

    "If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal, malicious, or extreme information. Training data is also at risk of being poisoned through tampering, error injection, or misleading actions by attackers. This can interfere with the model's probability distribution, reducing its accuracy and reliability."

    From AI Safety Governance Framework (TC2602024)

  17. "Eased production of and access to obscene, degrading, and/or abusive imagery which can cause harm, including synthetic child sexual abuse material (CSAM), and nonconsensual intimate images (NCII) of adults."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  18. 56.05.00 · Risk Category

    Harmful responses

    "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  19. 11.01.03 · Risk Sub-Category

    Representational Harms

    Erasing social groups

    people, attributes, or artifacts associated with specific social groups are systematically absent or under-represented... Design choices [143] and training data [212] influence which people and experiences are legible to an algorithmic system

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  20. 47.02.09 · Risk Sub-Category

    Ethical and social risks

    Bias and discrimination (value embedding)

    "Generative AI models may also be subject to the “value embedding” phenomenon.361 “Value embedding” refers to the fact that developers of generative AI models strive to minimize biased outputs by retraining their models based on normative values.362 Contemporary state-of- the-art models not only reflect the values embedded within their training data, they also undergo additional fine-tuning that follows a set of chosen rules and principles. Due to the absence of universally accepted standards, developers bear the responsibility of making decisions on sensitive issues. These practices lead to c

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  21. 52.03.02 · Risk Sub-Category

    Systemic Risks

    Ideological Homogenization from Value Embedding

    "The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world can make these value judgements unprecedently impactful, potentially leading to increased ideological homogenization."

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  22. 65.23.04 · Risk Sub-Category

    Non-technical risks (Societal impact)

    Impact on affected communities

    "It is important to include the perspectives or concerns of communities that are affected by model outcomes when designing and building models. Failing to include these perspectives makes it difficult to understand the relevant context for the model and to engender trust within these communities."

    From AI Risk Atlas (IBM2025)

  23. 02.07.01 · Risk Sub-Category

    Privacy Leakage

    Private Training Data

    "As recent LLMs continue to incorporate licensed, created, and publicly available data sources in their corpora, the potential to mix private data in the training corpora is significantly increased. The misused private data, also named as personally identifiable information (PII) [84], [86], could contain various types of sensitive data subjects, including an individual person’s name, email, phone number, address, education, and career. Generally, injecting PII into LLMs mainly occurs in two settings — the exploitation of web-collection data and the alignment with personal humanmachine convers

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  24. "Some of the broken systems discussed above are also very invasive of people’s privacy, controlling, for instance, the length of someone’s last romantic relationship [51]. More recently, ChatGPT was banned in Italy over privacy concerns and potential violation of the European Union’s (EU) General Data Protection Regulation (GDPR) [52]. The Italian data-protection authority said, “the app had experienced a data breach involving user conversations and payment information.” It also claimed that there was no legal basis to justify “the mass collection and storage of personal data for the purpose o

    From Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

  25. 06.02.00 · Risk Category

    Loss of privacy

    "AI offers the temptation to abuse someone's personal data, for instance to build a profile of them to target advertisements more effectively."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  26. "Face recognition technologies and their ilk pose significant privacy risks [47]. For example, we must consider certain ethical questions like: what data is stored, for how long, who owns the data that is stored, and can it be subpoenaed in legal cases [42]? We must also consider whether a human will be in the loop when decisions are made which rely on private data, such as in the case of loan decisions [37]."

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  27. 13.01.04 · Risk Sub-Category

    Impacts: The Technical Base System

    Privacy and Data Protection

    "Examining the ways in which generative AI systems providers leverage user data is critical to evaluating its impact. Protecting personal information and personal and group privacy depends largely on training data, training methods, and security measures."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  28. 19.04.02 · Risk Sub-Category

    Social AI Risks

    Privacy and safety concerns due to ubiquity of AI systems in economy and society (lack of social acceptance)

  29. 24.03.08 · Risk Sub-Category

    Malicious Uses

    Adversarial AI: Data and Model Exfiltration Attacks

    "Other forms of abuse can include privacy attacks that allow adversaries to exfiltrate or gain knowledge of the private training data set or other valuable assets. For example, privacy attacks such as membership inference can allow an attacker to infer the specific private medical records that were used to train a medical AI diagnosis assistant. Another risk of abuse centers around attacks that target the intellectual property of the AI assistant through model extraction and distillation attacks that exploit the tension between API access and confidentiality in ML models. Without the proper mi

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  30. 27.02.02 · Risk Sub-Category

    Instruction Attacks

    Prompt Leaking

    "By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  31. 31.03.00 · Risk Category

    Opaque Data Collection

    "When companies scrape personal information and use it to create generative AI tools, they undermine consumers' control of their personal information by using the information for a purpose for which the consumer did not consent."

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  32. 31.03.01 · Risk Sub-Category

    Opaque Data Collection

    Scraping to train data

    "When companies scrape personal information and use it to create generative AI tools, they undermine consumers’ control of their personal information by using the information for a purpose for which the consumer did not consent. The individual may not have even imagined their data could be used in the way the company intends when the person posted it online. Individual storing or hosting of scraped personal data may not always be harmful in a vacuum, but there are many risks. Multiple data sets can be combined in ways that cause harm: information that is not sensitive when spread across differ

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  33. 31.03.02 · Risk Sub-Category

    Opaque Data Collection

    Generative AI User Data

    Many generative AI tools require users to log in for access, and many retain user information, including contact information, IP address, and all the inputs and outputs or “conversations” the users are having within the app. These practices implicate a consent issue because generative AI tools use this data to further train the models, making their “free” product come at a cost of user data to train the tools. This dovetails with security, as mentioned in the next section, but best practices would include not requiring users to sign in to use the tool and not retaining or using the user-genera

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  34. "Vulnerable channel by which personal information may be accessed. The user may want their personal data to be kept private."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  35. 45.01.07 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of illegal collection and use of data)

    "The collection of AI training data and the interaction with users during service provision pose security risks, including collecting data without consent and improper use of data and personal information."

    From AI Safety Governance Framework (TC2602024)

  36. 45.01.10 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of data leakage)

    "In AI research, development, and applications, issues such as improper data processing, unauthorized access, malicious attacks, and deceptive interactions can lead to data and personal information leaks."

    From AI Safety Governance Framework (TC2602024)

  37. 45.02.03 · Risk Sub-Category

    Safety risks in AI Applications

    Cyberspace risks (Risks of information leakage due to improper usage)

    "Staff of government agencies and enterprises, if failing to use the AI service in a regulated and proper manner, may input internal data and industrial information into the AI model, leading to the leakage of work secrets, business secrets, and other sensitive business data."

    From AI Safety Governance Framework (TC2602024)

  38. 47.03.01 · Risk Sub-Category

    Legal challenges

    Privacy and data collection concerns (collecting personal information or personally identifiable information)

    "Generative AI developers train their models with extensive datasets often gathered through online web scraping of websites that may include personal data or personally identifiable information (PII). For most generative AI applications, such as initial model training, the primary concerns are the quantity, variety, and quality of the data, not whether they include personally identifiable information. However, some web-scraped datasets may inadvertently include personal data. Additionally, when downstream developers integrate generative AI into their products or services by fine- tuning a pre-

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  39. 65.05.02 · Risk Sub-Category

    Training Data Risks (Intellectual property)

    Confidential information in data

    "Confidential information might be included as part of the data that is used to train or tune the model."

    From AI Risk Atlas (IBM2025)

  40. 66.09.04 · Risk Sub-Category

    Privacy and Security

    Secondary use

    "The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, and (re)creating models from a public dataset."

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  41. 66.09.07 · Risk Sub-Category

    Privacy and Security

    Insecurity

    "carelessness in protecting collected personal data from leaks and improper access due to faulty data storage and data practices"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  42. 02.03.04 · Risk Sub-Category

    Unhelpful Uses

    Software Vulnerabilities

    "Programmers are accustomed to using code generation tools such as Github Copilot for program development, which may bury vulnerabilities in the program."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  43. 02.05.02 · Risk Sub-Category

    Hardware Vulnerabilities

    GPU Computation Platforms

    "The training of LLMs requires significant GPU resources, thereby introducing an additional security concern. GPU side-channel attacks have been developed to extract the parameters of trained models [159], [163]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  44. 02.05.03 · Risk Sub-Category

    Hardware Vulnerabilities

    Memory and Storage

    "Similar to conventional programs, hardware infrastructures can also introduce threats to LLMs. Memory-related vulnerabilities, such as rowhammer attacks [160], can be leveraged to manipulate the parameters of LLMs, giving rise to attacks such as the Deephammer attack [167], [168]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  45. 02.06.02 · Risk Sub-Category

    Issues on External Tools

    Exploiting External Tools for Attacks

    "Adversarial tool providers can embed malicious instructions in the APIs or prompts [84], leading LLMs to leak memorized sensitive information in the training data or users’ prompts (CVE2023-32786). As a result, LLMs lack control over the output, resulting in sensitive information being disclosed to external tool providers. Besides, attackers can easily manipulate public data to launch targeted attacks, generating specific malicious outputs according to user inputs. Furthermore, feeding the information from external tools into LLMs may lead to injection attacks [61]. For example, unverified in

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  46. 02.10.00 · Risk Category

    Model Attacks

    Model attacks exploit the vulnerabilities of LLMs, aiming to steal valuable information or lead to incorrect responses.

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  47. 02.10.01 · Risk Sub-Category

    Model Attacks

    Extraction Attacks

    "Extraction attacks [137] allow an adversary to query a black-box victim model and build a substitute model by training on the queries and responses. The substitute model could achieve almost the same performance as the victim model. While it is hard to fully replicate the capabilities of LLMs, adversaries could develop a domainspecific model that draws domain knowledge from LLMs"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  48. 02.10.02 · Risk Sub-Category

    Model Attacks

    Inference Attacks

    "Inference attacks [150] include membership inference attacks, property inference attacks, and data reconstruction attacks. These attacks allow an adversary to infer the composition or property information of the training data. Previous works [67] have demonstrated that inference attacks could easily work in earlier PLMs, implying that LLMs are also possible to be attacked"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  49. 02.10.03 · Risk Sub-Category

    Model Attacks

    Poisoning Attacks

    "Poisoning attacks [143] could influence the behavior of the model by making small changes to the training data. A number of efforts could even leverage data poisoning techniques to implant hidden triggers into models during the training process (i.e., backdoor attacks). Many kinds of triggers in text corpora (e.g., characters, words, sentences, and syntax) could be used by the attackers.""

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.