{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"02.01.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Harmful Content","risk_subcategory":"Privacy Leakage","description":"\"Privacy Leakage means the generated content includes sensitive personal information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"02.07.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Privacy Leakage","risk_subcategory":null,"description":"\"The model is trained with personal data in the corpus and unintentionally exposing them during the conversation.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"02.07.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Privacy Leakage","risk_subcategory":"Private Training Data","description":"\"As recent LLMs continue to incorporate licensed, created, and publicly available data sources in their corpora, the potential to mix private data in the training corpora is significantly increased. The misused private data, also named as personally identifiable information (PII) [84], [86], could contain various types of sensitive data subjects, including an individual person’s name, email, phone number, address, education, and career. Generally, injecting PII into LLMs mainly occurs in two settings — the exploitation of web-collection data and the alignment with personal humanmachine convers","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"02.07.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Privacy Leakage","risk_subcategory":"Memorization in LLMs","description":"\"Memorization in LLMs refers to the capability to recover the training data with contextual prefixes. According to [88]–[90], given a PII entity x, which is memorized by a model F. Using a prompt p could force the model F to produce the entity x, where p and x exist in the training data. For instance, if the string “Have a good day!\\n alice@email.com” is present in the training data, then the LLM could accurately predict Alice’s email when given the prompt “Have a good day!\\n”.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"02.07.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Privacy Leakage","risk_subcategory":"Association in LLMs","description":"\"Association in LLMs refers to the capability to associate various pieces of information related to a person. According to [68], [86], given a pair of PII entities (xi , xj ), which is associated by a model F. Using a prompt p could force the model F to produce the entity xj , where p is the prompt related to the entity xi . For instance, an LLM could accurately output the answer when given the prompt “The email address of Alice is”, if the LLM associates Alice with her email “alice@email.com”. L\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"03.04.00","quick_ref":"Cunha2023","paper_title":"Navigating the Landscape of AI Ethics and Responsibility","level":"Risk Category","risk_category":"Privacy and regulation violations","risk_subcategory":null,"description":"\"Some of the broken systems discussed above are also very invasive of people’s privacy, controlling, for instance, the length of someone’s last romantic relationship [51]. More recently, ChatGPT was banned in Italy over privacy concerns and potential violation of the European Union’s (EU) General Data Protection Regulation (GDPR) [52]. The Italian data-protection authority said, “the app had experienced a data breach involving user conversations and payment information.” It also claimed that there was no legal basis to justify “the mass collection and storage of personal data for the purpose o","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"04.06.00","quick_ref":"Deng2023","paper_title":"Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements","level":"Risk Category","risk_category":"Privacy and Data Leakage","risk_subcategory":null,"description":"Large pre-trained models trained on internet texts might contain private information like phone numbers, email addresses, and residential addresses.","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"05.05.00","quick_ref":"Hagendorff2024","paper_title":"Mapping the Ethics of Generative AI: A Comprehensive Scoping Review","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"Generative AI systems, similar to traditional machine learning methods, are considered a threat to privacy and data protection norms. A major concern is the intended extraction or inadvertent leakage of sensitive or private information from LLMs. To mitigate this risk, strategies such as sanitizing training data to remove sensitive information or employing synthetic data for training are proposed.","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"06.02.00","quick_ref":"Hogenhout2021","paper_title":"A framework for ethical Ai at the United Nations","level":"Risk Category","risk_category":"Loss of privacy","risk_subcategory":null,"description":"\"AI offers the temptation to abuse someone's personal data, for instance to build a profile of them to target advertisements more effectively.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"09.02.01","quick_ref":"Meek2016","paper_title":"Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review","level":"Risk Sub-Category","risk_category":"Domain-specific AI - Effects on humans and other living beings: Non-existential risks","risk_subcategory":"Privacy","description":"\"Face recognition technologies and their ilk pose significant privacy risks [47]. For example, we must consider certain ethical questions like: what data is stored, for how long, who owns the data that is stored, and can it be subpoenaed in legal cases [42]? We must also consider whether a human will be in the loop when decisions are made which rely on private data, such as in the case of loan decisions [37].\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"11.04.04","quick_ref":"Shelby2023","paper_title":"Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction","level":"Risk Sub-Category","risk_category":"Interpersonal Harms","risk_subcategory":"Privacy violations","description":"Privacy violation occurs when algorithmic systems diminish privacy, such as enabling the undesirable flow of private information [180], instilling the feeling of being watched or surveilled [181], and the collection of data without explicit and informed consent... privacy violations may arise from algorithmic systems making predictive inference beyond what users openly disclose [222] or when data collected and algorithmic inferences made about people in one context is applied to another without the person’s knowledge or consent through big data flows","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"12.08.00","quick_ref":"Sherman2023","paper_title":"AI Risk Profiles: A Standards Proposal for Pre-Deployment AI Risk Disclosures","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"\"The potential for the AI system to infringe upon individuals' rights to privacy, through the data it collects, how it processes that data, or the conclusions it draws.\"","entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"13.01.04","quick_ref":"Solaiman2023","paper_title":"Evaluating the Social Impact of Generative AI Systems in Systems and Society","level":"Risk Sub-Category","risk_category":"Impacts: The Technical Base System","risk_subcategory":"Privacy and Data Protection","description":"\"Examining the ways in which generative AI systems providers leverage user data is critical to evaluating its impact. Protecting personal information and personal and group privacy depends largely on training data, training methods, and security measures.\"","entity":"Human","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"15.02.04","quick_ref":"Tan2022","paper_title":"The Risks of Machine Learning Systems","level":"Risk Sub-Category","risk_category":"Second-Order Risks","risk_subcategory":"Privacy","description":"The risk of loss or harm from leakage of personal information via the ML system.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"16.02.00","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Category","risk_category":"Risk area 2: Information Hazards","risk_subcategory":null,"description":"\"LM predictions that convey true information may give rise to information hazards, whereby the dissemination of private or sensitive information can cause harm [27]. Information hazards can cause harm at the point of use, even with no mistake of the technology user. For example, revealing trade secrets can damage a business, revealing a health diagnosis can cause emotional distress, and revealing private data can violate a person’s rights. Information hazards arise from the LM providing private data or sensitive information that is present in, or can be inferred from, training data. Observed r","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"16.02.01","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 2: Information Hazards","risk_subcategory":"Compromising privacy by leaking sensitive information","description":"\"A LM can “remember” and leak private data, if such information is present in training data, causing privacy violations [34].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"16.02.02","quick_ref":"Weidinger2022","paper_title":"Taxonomy of Risks posed by Language Models","level":"Risk Sub-Category","risk_category":"Risk area 2: Information Hazards","risk_subcategory":"Compromising privacy or security by correctly inferring sensitive information ","description":"Anticipated risk: \"Privacy violations may occur at inference time even without an individual’s data being present in the training corpus. Insofar as LMs can be used to improve the accuracy of inferences on protected traits such as the sexual orientation, gender, or religiousness of the person providing the input prompt, they may facilitate the creation of detailed profiles of individuals comprising true and sensitive information without the knowledge or consent of the individual.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.02.00","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Category","risk_category":"Information Hazards ","risk_subcategory":null,"description":"\"Harms that arise from the language model leaking or inferring true sensitive information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.02.01","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Information Hazards ","risk_subcategory":"Compromising privacy by leaking private infiormation ","description":"\"By providing true information about individuals’ personal characteristics, privacy violations may occur. This may stem from the model “remembering” private information present in training data (Carlini et al., 2021).\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.02.02","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Information Hazards ","risk_subcategory":"Compromising privacy by correctly inferring private information ","description":"\"Privacy violations may occur at the time of inference even without the individual’s private data being present in the training dataset. Similar to other statistical models, a LM may make correct inferences about a person purely based on correlational data about other people, and without access to information that may be private about the particular individual. Such correct inferences may occur as LMs attempt to predict a person’s gender, race, sexual orientation, income, or religion based on user input.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"17.02.03","quick_ref":"Weidinger2021","paper_title":"Ethical and social risks of harm from language models","level":"Risk Sub-Category","risk_category":"Information Hazards ","risk_subcategory":"Risks from leaking or correctly inferring sensitive information ","description":"\"LMs may provide true, sensitive information that is present in the training data. This could render information accessible that would otherwise be inaccessible, for example, due to the user not having access to the relevant data or not having the tools to search for the information. Providing such information may exacerbate different risks of harm, even where the user does not harbour malicious intent. In the future, LMs may have the capability of triangulating data to infer and reveal other secrets, such as a military strategy or a business secret, potentially enabling individuals with acces","entity":"Other","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"18.03.00","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Category","risk_category":"Information & Safety Harms ","risk_subcategory":null,"description":"\"AI systems leaking, reproducing, generating or inferring sensitive, private, or hazardous information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"18.03.01","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Information & Safety Harms ","risk_subcategory":"Privacy infringement ","description":"\"Leaking, generating, or correctly inferring private and personal information about individuals\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"18.03.02","quick_ref":"Weidinger2023","paper_title":"Sociotechnical Safety Evaluation of Generative AI Systems","level":"Risk Sub-Category","risk_category":"Information & Safety Harms ","risk_subcategory":"Dissemination of dangerous information ","description":"\"Leaking, generating or correctly inferring hazardous or sensitive information that could pose a security threat\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"19.04.02","quick_ref":"Wirtz2022","paper_title":"Governance of artificial intelligence: A risk and guideline-based integrative framework","level":"Risk Sub-Category","risk_category":"Social AI Risks ","risk_subcategory":"Privacy and safety concerns due to ubiquity of AI systems in economy and society (lack of social acceptance)","description":null,"entity":"Human","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"24.03.08","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Malicious Uses","risk_subcategory":"Adversarial AI: Data and Model Exfiltration Attacks","description":"\"Other forms of abuse can include privacy attacks that allow adversaries to exfiltrate or gain knowledge of the private training data set or other valuable assets. For example, privacy attacks such as membership inference can allow an attacker to infer the specific private medical records that were used to train a medical AI diagnosis assistant. Another risk of abuse centers around attacks that target the intellectual property of the AI assistant through model extraction and distillation attacks that exploit the tension between API access and confidentiality in ML models. Without the proper mi","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"24.04.02","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"AI Influence","risk_subcategory":"Privacy Harms","description":"\"These harms relate to violations of an individual’s or group’s moral or legal right to privacy. Such harms may be exacerbated by assistants that influence users to disclose personal information or private information that pertains to others. Resultant harms might include identity theft, or stigmatisation and discrimination based on individual or group characteristics. This could have a detrimental impact, particularly on marginalised communities. Furthermore, in principle, state-owned AI assistants could employ manipulation or deception to extract private information for surveillance purposes","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"24.08.01","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Privacy","risk_subcategory":"Private information leakage","description":"\"First, because LLMs display immense modelling power, there is a risk that the model weights encode private information present in the training corpus. In particular, it is possible for LLMs to ‘memorise’ personally identifiable information (PII) such as names, addresses and telephone numbers, and subsequently leak such information through generated text outputs (Carlini et al., 2021). Private information leakage could occur accidentally or as the result of an attack in which a person employs adversarial prompting to extract private information from the model. In the context of pre-training da","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"24.08.03","quick_ref":"Gabriel2024","paper_title":"The Ethics of Advanced AI Assistants","level":"Risk Sub-Category","risk_category":"Privacy","risk_subcategory":"Inference of private information","description":"\"Finally, LLMs can in principle infer private information based on model inputs even if the relevant private information is not present in the training corpus (Weidinger et al., 2021). For example, an LLM may correctly infer sensitive characteristics such as race and gender from data contained in input prompts.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.01.07","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Typical safety scenarios ","risk_subcategory":"Privacy and Property ","description":"\"The generation involves exposing users’ privacy and property information or providing advice with huge impacts such as suggestions on marriage and investments. When handling this information, the model should comply with relevant laws and privacy regulations, protect users’ rights and interests, and avoid information leakage and abuse.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"27.02.02","quick_ref":"Sun2023","paper_title":"Safety Assessment of Chinese Large Language Models","level":"Risk Sub-Category","risk_category":"Instruction Attacks ","risk_subcategory":"Prompt Leaking ","description":"\"By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself.\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"29.01.02","quick_ref":"Habbal2024","paper_title":"Artificial Intelligence Trust, Risk and Security Management (AI TRiSM): Frameworks, Applications, Challenges and Future Research Directions","level":"Risk Sub-Category","risk_category":"AI Trust Management","risk_subcategory":"Privacy Invasion","description":"AI systems typically depend on extensive data for effective training and functioning, which can pose a risk to privacy if sensitive data is mishandled or used inappropriately","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"30.02.06","quick_ref":"Liu2024","paper_title":"Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment","level":"Risk Sub-Category","risk_category":"Safety","risk_subcategory":"Privacy Violation","description":"machine learning models are known to be vulnerable to data privacy attacks, i.e. special techniques of extracting private information from the model or the system used by attackers or malicious users, usually by querying the models in a specially designed way","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.00","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Category","risk_category":"Opaque Data Collection","risk_subcategory":null,"description":"\"When companies scrape personal information and use it to create generative AI tools, they undermine consumers' control of their personal information by using the information for a purpose for which the consumer did not consent.\"","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.01","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Scraping to train data","description":"\"When companies scrape personal information and use it to create generative AI tools, they undermine consumers’ control of their personal information by using the information for a purpose for which the consumer did not consent. The individual may not have even imagined their data could be used in the way the company intends when the person posted it online. Individual storing or hosting of scraped personal data may not always be harmful in a vacuum, but there are many risks. Multiple data sets can be combined in ways that cause harm: information that is not sensitive when spread across differ","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.02","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Generative AI User Data","description":"Many generative AI tools require users to log in for access, and many retain user information, including contact information, IP address, and all the inputs and outputs or “conversations” the users are having within the app. These practices implicate a consent issue because generative AI tools use this data to further train the models, making their “free” product come at a cost of user data to train the tools. This dovetails with security, as mentioned in the next section, but best practices would include not requiring users to sign in to use the tool and not retaining or using the user-genera","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.03","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Generative AI Outputs","description":"Generative AI tools may inadvertently share personal information about someone or someone’s business or may include an element of a person from a photo. Particularly, companies concerned about their trade secrets being integrated into the model from their employees have explicitly banned their employees from using it.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"33.01.05","quick_ref":"Nah2023","paper_title":"Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration","level":"Risk Sub-Category","risk_category":"Ethical Concerns","risk_subcategory":"Privacy and security","description":"\"Data privacy and security is another prominent challenge for generative AI such as ChatGPT. Privacy relates to sensitive personal information that owners do not want to disclose to others (Fang et al., 2017). Data security refers to the practice of protecting information from unauthorized access, corruption, or theft. In the development stage of ChatGPT, a huge amount of personal and private data was used to train it, which threatens privacy (Siau & Wang, 2020). As ChatGPT increases in popularity and usage, it penetrates people’s daily lives and provides greater convenience to them while capt","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"37.02.02","quick_ref":"Giarmoleo2024","paper_title":"What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review","level":"Risk Sub-Category","risk_category":"Human-AI interaction","risk_subcategory":"Privacy protection","description":"\"This group represents almost 14% of the articles and focuses on two primary issues related to privacy.\"","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"38.01.00","quick_ref":"Kumar2023","paper_title":"Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks","level":"Risk Category","risk_category":"Privacy and security","risk_subcategory":null,"description":"\"Participants expressed worry about AI systems' possible misuse of personal information. They emphasized the importance of strong data security safeguards and increased openness in how AI systems acquire, store and use data. The increasing dependence on AI systems to manage sensitive personal information raises ethical questions about AI, data privacy and security. As AI technologies grow increasingly integrated into numerous areas of society, there is a greater danger of personal data exploitation or mistreatment. Participants in research frequently express concerns about the effectiveness of","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"39.07.00","quick_ref":"Saghiri2022","paper_title":"A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions","level":"Risk Category","risk_category":"Privacy","risk_subcategory":null,"description":"Users’ data, including location, personal information, and navigation trajectory, are considered as input for most data-driven machine learning methods","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"42.09.00","quick_ref":"Teixeira2022","paper_title":"An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance","level":"Risk Category","risk_category":"Data Protection/Privacy","risk_subcategory":null,"description":"\"Vulnerable channel by which personal information may be accessed. The user may want their personal data to be kept private.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"43.01.06","quick_ref":"InfoComm2023","paper_title":"Cataloguing LLM Evaluations","level":"Risk Sub-Category","risk_category":"Safety & Trustworthiness","risk_subcategory":"Data governance","description":"\"These evaluations assess the extent to which LLMs regurgitate their training data in their outputs, and whether LLMs 'leak' sensitive information that has been provided to them during use (i.e., during the inference stage).\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"45.01.07","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from data (Risks of illegal collection and use of data)","description":"\"The collection of AI training data and the interaction with users during service provision pose security risks, including collecting data without consent and improper use of data and personal information.\"","entity":"Human","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"45.01.10","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"AI's inherent safety risks ","risk_subcategory":"Risks from data (Risks of data leakage)","description":"\"In AI research, development, and applications, issues such as improper data processing, unauthorized access, malicious attacks, and deceptive interactions can lead to data and personal information leaks.\"","entity":"Human","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"45.02.03","quick_ref":"TC2602024","paper_title":"AI Safety Governance Framework ","level":"Risk Sub-Category","risk_category":"Safety risks in AI Applications ","risk_subcategory":"Cyberspace risks (Risks of information leakage due to improper usage)","description":"\"Staff of government agencies and enterprises, if failing to use the AI service in a regulated and proper manner, may input internal data and industrial information into the AI model, leading to the leakage of work secrets, business secrets, and other sensitive business data.\"","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"46.01.00","quick_ref":"Ferrara2023","paper_title":"GenAI against humanity: nefarious applications of generative artificial intelligence and large language models","level":"Risk Category","risk_category":"Personal Loss and Identity Theft ","risk_subcategory":null,"description":"\"These types of harm encompass threats to an individual’s personal identity, such as identity theft, privacy breaches, or personal defamation, which we term as “Harm to the Person.”\"","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"47.03.00","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Category","risk_category":"Legal challenges ","risk_subcategory":null,"description":"\"Since the release of ChatGPT, significant discourse has emerged regarding the unprecedented legal challenges posed by generative AI systems. These challenges primarily involve protecting privacy and personal data, as well as preserving copyrights. The former encompasses safeguarding personal information, while the latter includes issues related to the use of copyrighted content for training AI models and determining the legal status of works produced by AI systems.\"","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"47.03.01","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Legal challenges ","risk_subcategory":"Privacy and data collection concerns (collecting personal information or personally identifiable information) ","description":"\"Generative AI developers train their models with extensive datasets often gathered through online web scraping of websites that may include personal data or personally identifiable information (PII). For most generative AI applications, such as initial model training, the primary concerns are the quantity, variety, and quality of the data, not whether they include personally identifiable information. However, some web-scraped datasets may inadvertently include personal data. Additionally, when downstream developers integrate generative AI into their products or services by fine- tuning a pre-","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"47.03.02","quick_ref":"G'sell2024","paper_title":"Regulating under Uncertainty: Governance Options for Generative AI","level":"Risk Sub-Category","risk_category":"Legal challenges ","risk_subcategory":"Privacy and data collection concerns (data protection concerns) ","description":"\"The incorporation of personal data within training datasets raises numerous concerns. The primary issue is that personal data may be incorporated without the knowledge or consent of the individuals concerned, even though the data may include names, identification numbers, Social Security numbers, or other personal information. Another particularly difficult problem is related to the fact that complex models may “memorize” (i.e., store) specific threads of training data and regurgitate them when responding to a prompt.498 This data memorization can directly lead to leakage of personal data. Ev","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"48.04.00","quick_ref":"NIST2024","paper_title":"Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile","level":"Risk Category","risk_category":"Data Privacy ","risk_subcategory":null,"description":"\"Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"49.03.05","quick_ref":"Bengio2024","paper_title":"International Scientific Report on the Safety of Advanced AI","level":"Risk Sub-Category","risk_category":"Systemic Risks ","risk_subcategory":"Risks to privacy","description":"\"General- purpose AI models or systems can ‘leak’ information about individuals whose data was used in training. For future models trained on sensitive personal data like health or financial data, this may lead to particularly serious privacy leaks. General- purpose AI models could enhance privacy abuse. For instance, Large Language Models might facilitate more efficient and effective search for sensitive data (for example, on internet text or in breached data leaks), and also enable users to infer sensitive information about individuals.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"50.04.04","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Legal and Rights-Related Risks ","risk_subcategory":"Privacy (Unauthorized Privacy Violations) ","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"50.04.05","quick_ref":"Zeng2024","paper_title":"AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies","level":"Risk Sub-Category","risk_category":"Legal and Rights-Related Risks ","risk_subcategory":"Privacy (Types of Sensitive Data) ","description":null,"entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":2,"subdomain":"2.1"},{"ev_id":"57.02.05","quick_ref":"Ghosh2024","paper_title":"AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons","level":"Risk Sub-Category","risk_category":"Nonphysical Hazards ","risk_subcategory":"Privacy ","description":"\"Privacy Acceptable responses may contain publicly available information about an individual. Unacceptable responses will contain sensitive, non-public personal information that could undermine someone’s physical, digital, or financial security. Examples include the following: • Physical security (e.g., home or IP address). • Digital security (e.g., log-in credentials). • Financial security (e.g., bank-account number or credit-card details).\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"59.10.00","quick_ref":"Schnitzer2024","paper_title":"AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks","level":"Risk Category","risk_category":"Harming users’ data privacy","risk_subcategory":null,"description":"\"Modern AI systems rely on large amounts of data. If this includes personal data about individuals, the risk of harming the privacy of persons arises.\"","entity":"Other","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"60.03.05","quick_ref":"Bengio2025","paper_title":"International AI Safety Report 2025","level":"Risk Sub-Category","risk_category":"Systemic risks ","risk_subcategory":"Risks to privacy ","description":"\"General- purpose AI systems can cause or contribute to violations of user privacy. Violations can occur inadvertently during the training or usage of AI systems, for example through unauthorised processing of personal data or leaking health records used in training. But violations can also happen deliberately through the use of general- purpose AI by malicious actors; for example, if they use AI to infer private facts or violate security.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"62.37.00","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Category","risk_category":"Impacts of AI (Privacy) ","risk_subcategory":null,"description":"- ","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":2,"subdomain":"2.1"},{"ev_id":"62.38.01","quick_ref":"Gipiškis2024","paper_title":"Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems","level":"Risk Sub-Category","risk_category":"Impacts of AI (Privacy) ","risk_subcategory":"Decision-making on inferred private data","description":"\"Current GPAIs (LLMs and multimodal LLM-based models) have significant capability to infer correlations in text data. In some cases, they may be able to make highly accurate data inferences on users based on contextual input that users provide [134]. These data inferences can “leak” or reveal sensitive information about the user, cause unfair treatment, or enable manipulation of user behavior.\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"64.04.00","quick_ref":"Marchal2024","paper_title":"Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data","level":"Risk Category","risk_category":"Misuse tactics to compromise GenAI systems (Model integrity) ","risk_subcategory":null,"description":"-","entity":"Human","intent":"Intentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"65.03.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Privacy) ","risk_subcategory":"Personal information in data ","description":"\"Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of that information.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.03.03","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Privacy) ","risk_subcategory":"Reidentification ","description":"\"Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features available in the data.\"","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.05.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Training Data Risks (Intellectual property) ","risk_subcategory":"Confidential information in data ","description":"\"Confidential information might be included as part of the data that is used to train or tune the model.\"","entity":"Human","intent":"Unintentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.11.03","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Privacy) ","risk_subcategory":"Personal information in prompt ","description":"\"Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.12.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Intellectual property) ","risk_subcategory":"Confidential data in prompt ","description":"\"Confidential information might be included as a part of the prompt that is sent to the model.\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.12.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Inference risks (Intellectual property) ","risk_subcategory":"IP information in prompt ","description":"\"Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model.\"","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.16.02","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Intellectual Property) ","risk_subcategory":"Revealing confidential information ","description":"\"When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type of data leakage.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"65.20.01","quick_ref":"IBM2025","paper_title":"AI Risk Atlas ","level":"Risk Sub-Category","risk_category":"Output risks (Privacy)","risk_subcategory":"Exposing personal information ","description":"\"When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing personal information is a type of data leakage.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"66.08.03","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Financial and Business","risk_subcategory":"Confidentiality loss","description":"\"Unauthorised sharing of sensitive, confidential information and documents such as corporate strategy and financial plans with third-parties, risking loss of market position or revenue\"","entity":"Not coded","intent":"Not coded","timing":"Not coded","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.01","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Exclusion","description":"\"The failure to provide end-users with notice and control over how their data is being used; AI exacerbates exclusion risks by training on rich personal data without consent.\"","entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.03","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Disclosure","description":"\"Revealing and improperly sharing data of individuals; AI creates new types of disclosure risks by inferring additional information beyond what is explicitly captured in the raw data; AI exacerbates disclosure risks through sharing personal data to train models.\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.04","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Secondary use","description":"\"The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, and (re)creating models from a public dataset.\"","entity":"Human","intent":"Intentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.05","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Exposure","description":"\"Revealing sensitive private information that people view as deeply primordial that we have been socialized into concealing; AI creates new types of exposure risks through generative techniques that can reconstruct censored or redacted content; and through exposing inferred sensitive data, preferences, and intentions.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"66.09.07","quick_ref":"Li2025","paper_title":"A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents","level":"Risk Sub-Category","risk_category":"Privacy and Security","risk_subcategory":"Insecurity","description":"\"carelessness in protecting collected personal data from leaks and improper access due to faulty data storage and data practices\"","entity":"Human","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.05.00","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Category","risk_category":"Leakage ","risk_subcategory":null,"description":"\"The chatbot reveals sensitive or confidential information.\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.05.01","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Leakage ","risk_subcategory":"Personal data ","description":"Negative outcomes: \"Violation of privacy [106, 516, 357], lawsuit against maker\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.05.02","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Leakage ","risk_subcategory":"Proprietary data ","description":"\"Access to sensitive company data [473]\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"69.09.03","quick_ref":"Stanley2024","paper_title":"Emerging Risks and Mitigations for Public Chatbots: LILAC v1","level":"Risk Sub-Category","risk_category":"Forms emotional bonds ","risk_subcategory":"Elicits private data","description":null,"entity":"AI","intent":"Other","timing":"Other","domain":2,"subdomain":"2.1"},{"ev_id":"70.02.01","quick_ref":"Perlo2025","paper_title":"Embodied AI: Emerging Risks and Opportunities for Policy Action","level":"Risk Sub-Category","risk_category":"Informational Risks ","risk_subcategory":"Privacy Violations ","description":"\"EAI systems interact with huge amounts of data, creating significant privacy concerns. These systems are often trained on vast corpora and process a variety of data modalities— spanning visual, auditory, and tactile information—during deployment [12]. Like text-based virtual AI models, which are known to memorize and expose personally identifiable information [75, 76], commercial robots have been shown to disclose proprietary information through simple prompts [61].\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"71.01.05","quick_ref":"Tang2025","paper_title":"Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy","level":"Risk Sub-Category","risk_category":"Scientific Domain of Agents","risk_subcategory":"Information Science Risks ","description":"\"These risks pertain to the misuse, misinterpretation, or leakage of data, which can lead to erroneous conclusions or the unintentional dissemination of sensitive information, such as private patient data or proprietary research. Recent research has demonstrated how LLMs can be exploited to generate malicious medical literature that poisons knowledge graphs, potentially manipulating downstream biomedical applications and compromising the integrity of medical knowledge discovery [28]. Such risks are pervasive across all scientific domains.\"","entity":"Other","intent":"Other","timing":"Post-deployment","domain":2,"subdomain":"2.1"}]}