MIT AI Risk Repository · Risk Category · 05.14.00

Evaluation - Auditing

Description

Closely related to other clusters like AI safety, fairness, or harmful content, papers stress the importance of evaluating generative AI systems both in a narrow technical way as well as in a broader sociotechnical impact assessment focusing on pre-release audits as well as post-deployment monitoring. Ideally, these evaluations should be conducted by independent third parties. In terms of technical LLM or text-to-image model audits, papers furthermore criticize a lack of safety benchmarking for languages other than English.

From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Domain
Subdomain
Causal entity
Not coded
Intent
Not coded
Timing
Not coded

Other entries from Hagendorff2024