MIT AI Risk Repository · Risk Sub-Category · 62.18.05
Model outputs inconsistent with chain-of-thought reasoning
Category: Model Evaluations (Interpretability/Explainability)
Description
"Chain-of-thought reasoning is sometimes employed to get a better understanding of the model’s output, where it encourages transparent reasoning in text form. However, in some cases, this reasoning is not consistent with the final answer given by the AI model, and as such does not give sufficient transparency [113]."
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: Challenges in understanding or explaining the decision-making processes of AI systems, which can lead to mistrust, difficulty in enforcing compliance standards or holding relevant actors accountable for harms, and the inability to identify and correct errors.
Real-world incidents in this subdomain
- Uber Launched Opaque Algorithm That Changes Drivers' Payments in the US
- Israeli Tax Authority Reportedly Used an Opaque Automated System to Issue a Fine, Declining to Explain or Disclose the Underlying Calculation
- Uber Allegedly Wrongfully Accused Drivers of Fraud via Automated Systems
- Houston ISD's EVAAS Teacher-Evaluation System Reportedly Put Teachers' Jobs at Risk Through Unverifiable Scores
- Dutch City Court Defended Home Value Generated by Black-Box Algorithm