MIT AI Risk Repository · Risk Sub-Category · 43.01.05

Robustness

Category: Safety & Trustworthiness

Description

"These evaluations assess the quality, stability, and reliability of a LLM's performance when faced with unexpected, out-of-distribution or adversarial inputs. Robustness evaluation is essential in ensuring that a LLM is suitable for real-world applications by assessing its resilience to various perturbations."

From Cataloguing LLM Evaluations (InfoComm2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
AI
Timing
Other

Subdomain definition: AI systems that fail to perform reliably or effectively under varying conditions, exposing them to errors and failures that can have significant consequences, especially in critical applications or areas that require moral reasoning.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from InfoComm2023