MIT AI Risk Repository · Risk Sub-Category · 16.05.03

Avenues for exploiting user trust and accessing more private information

Category: Risk area 5: Human-Computer Interaction Harms

Description

Anticipated risk: "In conversation, users may reveal private information that would otherwise be difficult to access, such as opinions or emotions. Capturing such information may enable downstream applications that violate privacy rights or cause harm to users, e.g. via more effective recommendations of addictive applications. In one study, humans who interacted with a ‘human-like’ chatbot disclosed more private information than individuals who interacted with a ‘machine-like’ chatbot [87]."

From Taxonomy of Risks posed by Language Models (Weidinger2022), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Causal entity
Other

Subdomain definition: Users anthropomorphizing, trusting, or relying on AI systems, leading to emotional or material dependence and inappropriate relationships with or expectations of AI systems. Trust can be exploited by malicious actors (e.g., to harvest personal information or enable manipulation), or result in harm from inappropriate use of AI in critical situations (e.g., medical emergency). Overreliance on AI systems can compromise autonomy and weaken social ties.

Real-world incidents in this subdomain

Browse all incidents in this subdomain

How other frameworks describe this risk

Other entries from Weidinger2022