MIT AI Risk Repository · Risk Sub-Category · 62.18.02
Misunderstanding or overestimating the results and scope of interpretability techniques
Category: Model Evaluations (Interpretability/Explainability)
Description
"The results of explainability techniques are not free of bias and require careful interpretation. Users might develop a false sense of security or reliability if the resulting explanations align with their initial beliefs, leading to confirmation bias and an overestimation of abilities of these techniques [24]."
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).