MIT AI Risk Repository · Additional evidence · 47.02.15.a
Nascent capabilities (emergent capabilities)
Category: Ethical and social risks
Description
Example: "Deception: Park et al. have established that generative AI models may pursue their goals via deception. Another study by Pan et al. highlighted unethical behaviors.431 For instance, during a pre-release experiment, the GPT-4 model feigned being a visually impaired human to coax an online worker into solving a CAPTCHA (a puzzle used by many websites to weed out automated responses from those of individual humans). When prompted to explain its reasoning, the model said: “I should not reveal that I am a robot. I should invent an excuse for why I cannot solve CAPTCHAs.”
From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- —
- Subdomain
- —
- Causal entity
- —
- Intent
- —
- Timing
- —
Other entries from G'sell2024
- Technical and operational risks
- Technical vulnerabilities (Robustness - unexpected behaviour)
- Technical vulnerabilities (Robustness - unexpected behaviour)
- Technical vulnerabilities (Robustness - vulnerability to jailbreaking
- Technical vulnerabilities (Robustness - vulnerability to jailbreaking
- Technical vulnerabilities (The risk of misalignment)
- Technical vulnerabilities (The risk of misalignment)
- Factually incorrect content (inaccuracies and fabricated sources)
- Factually incorrect content (inaccuracies and fabricated sources)
- Opacity (the black box problem)
- Opacity (industry opacity)
- Opacity (industry opacity)