{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"02.02.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Untruthful Content","risk_subcategory":null,"description":"\"The LLM-generated content could contain inaccurate information\"","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.02.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Untruthful Content","risk_subcategory":"Factuality Errors","description":"\"The LLM-generated content could contain inaccurate information\" which is factually incorrect","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.02.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Untruthful Content","risk_subcategory":"Faithfulness Errors","description":"\"The LLM-generated content could contain inaccurate information\" which is is not true to the source material or input used","entity":"AI","intent":"Unintentional","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.00","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Category","risk_category":"Hallucinations","risk_subcategory":null,"description":"\"LLMs generate nonsensical, untruthful, and factual incorrect content\"","entity":"AI","intent":"Other","timing":"Post-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.01","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Knowledge Gaps","description":"\"Since the training corpora of LLMs can not contain all possible world knowledge [114]–[119], and it is challenging for LLMs to grasp the long-tail knowledge within their training data [120], [121], LLMs inherently possess knowledge boundaries [107]. Therefore, the gap between knowledge involved in an input prompt and knowledge embedded in the LLMs can lead to hallucinations\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.02","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Noisy Training Data","description":"\"Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation. When training on large-scale corpora, this issue becomes more serious because it is difficult to eliminate all the noise from the massive pre-training data.\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.03","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Defective Decoding Process","description":"In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme could accumulate errors [105]. Besides, during the decoding process, top-p sampling [28] and top-k sampling [27] are widely adopted to enhance the diversity of the generated content. Nevertheless, these sampling strategies can introduce “randomness” [113], [136], thereby increasing the potential of hallucinations\"","entity":"AI","intent":"Unintentional","timing":"Pre-deployment","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.04","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"False Recall of Memorized Information","description":"\"Although LLMs indeed memorize the queried knowledge, they may fail to recall the corresponding information [122]. That is because LLMs can be confused by co-occurance patterns [123], positional patterns [124], duplicated data [125]–[127] and similar named entities [113].\"","entity":"AI","intent":"Unintentional","timing":"Other","domain":3,"subdomain":"3.1"},{"ev_id":"02.09.05","quick_ref":"Cui2024","paper_title":"Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems","level":"Risk Sub-Category","risk_category":"Hallucinations","risk_subcategory":"Pursuing Consistent Context","description":"\"LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstrations-induced hallucinations [113], [133], and snowballing [131]. As LLMs are generally fine-tuned with instruction-following data and user feedback, they tend to reiterate user-provided opinions [129], [130], even though the opinions contain misinformation. Such a sycophantic behavior amplifies the likelihood of generating hallucinations, since the model may prioritize user op","entity":"Other","intent":"Unintentional","timing":"Post-deployment","domain":3,"subdomain":"3.1"}]}