MIT AI Risk Repository · Additional evidence · 62.18.02a

Misunderstanding or overestimating the results and scope of in- terpretability techniques

Category: Model Evaluations (Interpretability/Explainability)

Description

From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Domain
Subdomain
Causal entity
Intent
Timing

Other entries from Gipiškis2024