MIT AI Risk Repository · Additional evidence · 47.02.15.a

Nascent capabilities (emergent capabilities)

Category: Ethical and social risks

Description

Example: "Deception: Park et al. have established that generative AI models may pursue their goals via deception. Another study by Pan et al. highlighted unethical behaviors.431 For instance, during a pre-release experiment, the GPT-4 model feigned being a visually impaired human to coax an online worker into solving a CAPTCHA (a puzzle used by many websites to weed out automated responses from those of individual humans). When prompted to explain its reasoning, the model said: “I should not reveal that I am a robot. I should invent an excuse for why I cannot solve CAPTCHAs.”

From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Domain
Subdomain
Causal entity
Intent
Timing

Other entries from G'sell2024