The OpenAI chatbot platform ChatGPT has introduced two new security features: Lockdown Mode and Elevated Risk labels. These features aim to mitigate the risk of prompt injection attacks, which can lead to sensitive data exfiltration when users interact with external systems through ChatGPT.
What Happened
Lockdown Mode is an optional advanced security setting that constrains how ChatGPT interacts with external systems. When enabled, it deterministically disables certain tools and capabilities that an attacker could attempt to exploit through users' conversations or connected apps. For example, live network requests are blocked, and browsing is limited to cached content.
Elevated Risk labels provide in-product guidance for features that may introduce additional security risk when connecting AI products with apps and the web. These labels explain what a feature does, what changes when it's enabled, what risks may be introduced, and when its use is appropriate.
Background and Context
Prompt injection attacks involve embedding hidden instructions in external content that ChatGPT reads, which can lead to sensitive data exfiltration. This type of attack requires three conditions to be true: the AI system has access to private data, it's exposed to untrusted content, and it has an outbound data channel. Lockdown Mode solves this problem by eliminating the third leg – the outbound channel.
OpenAI has identified two patterns that are particularly vulnerable to prompt injection attacks: professionals who paste confidential client documents into ChatGPT and use Deep Research or Agent Mode to cross-reference them against external sources, and those who use ChatGPT in basic chat mode for everyday writing help but still interact with external systems.
Why It Matters to the Industry
The introduction of Lockdown Mode and Elevated Risk labels is significant for adult-industry platforms and operators. These features can help mitigate the risk of prompt injection attacks, which can lead to sensitive data exfiltration. This is particularly important in industries where sensitive information is shared, such as financial reports or HR documents.
Lockdown Mode's ability to constrain how ChatGPT interacts with external systems can also be beneficial for adult-industry platforms that require high levels of security and control over user interactions. Elevated Risk labels provide transparency into the potential risks associated with certain features, allowing users to make informed decisions about their usage.
What Comes Next
OpenAI plans to expand availability of Lockdown Mode to consumer users in the future. The company will also continue to update which features carry Elevated Risk labels over time to best communicate risk to users. This suggests that OpenAI is committed to strengthening its safety and security safeguards, especially for novel, emerging, or growing risks.
Key Facts
- Lockdown Mode is an optional advanced security setting in ChatGPT that constrains how the platform interacts with external systems.
- Elevated Risk labels provide in-product guidance for features that may introduce additional security risk when connecting AI products with apps and the web.
- Prompt injection attacks involve embedding hidden instructions in external content that ChatGPT reads, which can lead to sensitive data exfiltration.
- Lockdown Mode solves the problem of prompt injection attacks by eliminating the outbound channel.
- Elevated Risk labels apply to features across ChatGPT, ChatGPT Atlas, and Codex.
OpenAI's introduction of Lockdown Mode and Elevated Risk labels is a significant step towards strengthening its safety and security safeguards. These features have the potential to mitigate the risk of prompt injection attacks and provide transparency into potential risks associated with certain features. As the adult industry continues to rely on AI-powered platforms, it's essential that these technologies are developed with security and control in mind.