MIT AI Risk Repository · Risk Sub-Category · 47.01.01

Technical vulnerabilities (Robustness - unexpected behaviour)

Category: Technical and operational risks

Description

"There is no assurance that generative AI models will consistently behave as their developers and users intend. Unwanted content is not necessarily due to intentional adversarial behavior. Generative AI models can unexpectedly produce potentially harmful content, including materials that are racist, discriminatory, or sexually explicit, or that promote violence, terrorism, or hate."

From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI
Intent
Other

Subdomain definition: AI systems that fail to perform reliably or effectively under varying conditions, exposing them to errors and failures that can have significant consequences, especially in critical applications or areas that require moral reasoning.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from G'sell2024