MIT AI Risk Repository · Risk Category · 05.14.00
Evaluation - Auditing
Description
Closely related to other clusters like AI safety, fairness, or harmful content, papers stress the importance of evaluating generative AI systems both in a narrow technical way as well as in a broader sociotechnical impact assessment focusing on pre-release audits as well as post-deployment monitoring. Ideally, these evaluations should be conducted by independent third parties. In terms of technical LLM or text-to-image model audits, papers furthermore criticize a lack of safety benchmarking for languages other than English.
From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).