{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-12"}
{"rows":[{"ev_id":"31.03.00","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Category","risk_category":"Opaque Data Collection","risk_subcategory":null,"description":"\"When companies scrape personal information and use it to create generative AI tools, they undermine consumers' control of their personal information by using the information for a purpose for which the consumer did not consent.\"","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.01","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Scraping to train data","description":"\"When companies scrape personal information and use it to create generative AI tools, they undermine consumers’ control of their personal information by using the information for a purpose for which the consumer did not consent. The individual may not have even imagined their data could be used in the way the company intends when the person posted it online. Individual storing or hosting of scraped personal data may not always be harmful in a vacuum, but there are many risks. Multiple data sets can be combined in ways that cause harm: information that is not sensitive when spread across differ","entity":"Human","intent":"Intentional","timing":"Pre-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.02","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Generative AI User Data","description":"Many generative AI tools require users to log in for access, and many retain user information, including contact information, IP address, and all the inputs and outputs or “conversations” the users are having within the app. These practices implicate a consent issue because generative AI tools use this data to further train the models, making their “free” product come at a cost of user data to train the tools. This dovetails with security, as mentioned in the next section, but best practices would include not requiring users to sign in to use the tool and not retaining or using the user-genera","entity":"Human","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"},{"ev_id":"31.03.03","quick_ref":"EPIC2023","paper_title":"Generating Harms - Generative AI's impact and paths forwards","level":"Risk Sub-Category","risk_category":"Opaque Data Collection","risk_subcategory":"Generative AI Outputs","description":"Generative AI tools may inadvertently share personal information about someone or someone’s business or may include an element of a person from a photo. Particularly, companies concerned about their trade secrets being integrated into the model from their employees have explicitly banned their employees from using it.","entity":"AI","intent":"Unintentional","timing":"Post-deployment","domain":2,"subdomain":"2.1"}]}