MIT AI Risk Repository · Risk Sub-Category · 62.16.06
General Evaluations (Biased evaluations of encoded human values)
Category: Model Evaluations
Description
"Encoded human values in AI models that are easier to evaluate might be preferred for inclusion in evaluations over those that are more difficult to measure [13]. This might come at the expense of more desirable but harder-to-quantify values. This bias can lead to an imbalance, where easier-to-measure values dominate the evaluation process, while other important values are underrepresented."
From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 6.5 Governance failure
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Pre-deployment
Subdomain definition: Inadequate regulatory frameworks and oversight mechanisms failing to keep pace with AI development, leading to ineffective governance and the inability to manage AI risks appropriately.
Real-world incidents in this subdomain
- Nippon Life Alleged ChatGPT Practiced Law Without a License in Illinois Disability Case
- OpenAI Allegedly Did Not Alert RCMP After ChatGPT Flagged Violent Chats Before British Columbia School Shooting
- Remotely Operated Taser-Armed Drones Proposed by Taser Manufacturer as Defense for School Shootings in the US
How other frameworks describe this risk
- Mobility
- Liability issues in case of accidents
- Faster scientific progress makes it harder for governance to keep pace with development
- Type 1: Diffusion of responsibility
- Products Liability Law
- Institutional responsibilities
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
- Governance - Regulation