AI incident #473 ·
Bing Chat's Initial Prompts Revealed by Early Testers Through Prompt Injection
What happened
Early testers of Bing Chat successfully used prompt injection to reveal its built-in initial instructions, which contains a list of statements governing ChatGPT's interaction with users.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (1)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged deployer
- Microsoft
- Alleged developer
- Microsoft, Openai
- Alleged harmed party
- Microsoft
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Privacy & Security
- Risk subdomain
- 2.2 AI system security vulnerabilities and attacks
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 2.2.
- Privacy loss
"Privacy loss - Unwarranted exposure of an individual’s private life or personal data through cyberattacks, doxxing, etc."
- Exploiting Limited Generalization of Safety Finetuning
"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. usi...
- Jailbreaks and Prompt Injections Threaten Security of LLMs
"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardiz...
- “Model Psychology” Attacks
"LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andr...
- Attacking LLMs via Additional Modalities a
"LLMs can now process modalities other than text, e.g. images or video frames (OpenAI, 2023c; Gemini Team, 2023). Several studies show that gradient-based attacks on multimodal models are easy and effective (Carlini et a...
- Vulnerability to Poisoning and Backdoors
"The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other...
- Model Attacks
Model attacks exploit the vulnerabilities of LLMs, aiming to steal valuable information or lead to incorrect responses.
- Hardware Vulnerabilities
"The vulnerabilities of hardware systems for training and inferencing brings issues to LLM-based applications."
Incidents in the same risk subdomain
- COEMPT Quality Assurance Engineers Allegedly Violated Indian CBSE Student Data Privacy Rights by Processing It with Google Gemini
- Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu
- Meta Internal AI Agent Reportedly Gave Advice That Allegedly Exposed Sensitive Data to Unauthorized Employees
- CodeWall's Autonomous Agent Reportedly Obtained Unauthorized Access to McKinsey's Lilli AI Platform Database
- Anthropic Said DeepSeek, Moonshot, and MiniMax Used Fraudulent Accounts and Proxies to Illicitly Distill Claude Capabilities at Scale
- DJI Romo Cloud Authorization Bug Reportedly Exposed Camera, Microphone, and Home-Mapping Data From Nearly 7,000 Robot Vacuums
Other incidents involving Microsoft
- Eightfold AI Hiring Tools Allegedly Secretly Scored Job Applicants for Employers
- Microsoft's Windows Recall Allegedly Stores Passwords and Social Security Numbers in Preview Mode
- Microsoft 365 Copilot Vulnerability Allegedly Allowed File Access Without Audit Log Entry
- Microsoft Copilot Reportedly Able to Access Cached Data from Since-Private GitHub Repositories
- Microsoft Copilot Designer Reportedly Generated Inappropriate AI Images
- Microsoft AI Poll Allegedly Causes Reputational Harm of The Guardian Newspaper