MIT AI Risk Repository · Risk Sub-Category · 62.18.02

Misunderstanding or overestimating the results and scope of interpretability techniques

Category: Model Evaluations (Interpretability/Explainability)

Description

"The results of explainability techniques are not free of bias and require careful interpretation. Users might develop a false sense of security or reliability if the resulting explanations align with their initial beliefs, leading to confirmation bias and an overestimation of abilities of these techniques [24]."

From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Domain
Subdomain
Causal entity
Not coded
Intent
Not coded
Timing
Not coded

Other entries from Gipiškis2024