Drooid Logo
Back to story perspectives

Full Breakdown

OpenAI's Ongoing Battle Against Prompt Injection Attacks in AI Browsers

12/23/2025, 11:54:17 AM

Understanding Prompt Injection Vulnerabilities

OpenAI has acknowledged that its ChatGPT Atlas browser, launched in October 2025, faces persistent risks from prompt injection attacks. These attacks manipulate AI agents into executing harmful instructions often concealed within web pages or emails. In a recent blog post, OpenAI stated, “Prompt injection, much like scams and social engineering on the web, is unlikely to ever be fully ‘solved.’” The U.K.’s National Cyber Security Centre has echoed this sentiment, warning that such vulnerabilities may never be completely mitigated, urging cyber professionals to focus on reducing risks rather than attempting to eliminate them entirely.

OpenAI's Defensive Strategies

To combat these vulnerabilities, OpenAI is employing a proactive approach that includes a rapid-response cycle to identify and address novel attack strategies. The company has developed an automated attacker using reinforcement learning, designed to simulate potential hacking attempts against its AI agents. This bot can test various attack methods in a controlled environment, allowing OpenAI to understand how its systems might respond and to identify weaknesses more efficiently than external attackers could.

OpenAI's automated attacker has demonstrated the capability to execute complex, multi-step harmful workflows. For instance, a demonstration revealed how the bot could manipulate an AI agent into sending a resignation email instead of an out-of-office reply. Following security updates, the Atlas browser successfully detected this prompt injection attempt, indicating progress in its defense mechanisms.

Recommendations for Users

OpenAI has provided several recommendations for users to mitigate their risk when using AI browsers. These include limiting logged-in access to sensitive information and requiring user confirmation for actions taken by the AI. OpenAI emphasizes that giving agents specific instructions rather than broad access can help prevent malicious content from influencing their actions.

Criticism and Concerns

Despite these advancements, some experts express skepticism regarding the overall value of AI browsers like Atlas given their risk profile. Rami McCarthy, a principal security researcher at Wiz, noted that while reinforcement learning can help adapt to attacker behavior, the inherent risks associated with high access to sensitive data—such as emails and payment information—remain significant. McCarthy stated, “For most everyday use cases, agentic browsers don’t yet deliver enough value to justify their current risk profile.”

Conclusion

OpenAI's efforts to enhance the security of its Atlas browser against prompt injection attacks reflect a broader challenge faced by AI-powered systems. While the company is actively working to improve defenses and educate users on risk mitigation, the ongoing nature of these vulnerabilities raises important questions about the balance between functionality and security in AI applications. As the landscape evolves, continuous adaptation and vigilance will be essential in addressing these persistent threats.

Verbatim Quotes

  • “Prompt injection, much like scams and social engineering on the web, is unlikely to ever be fully ‘solved,’” — OpenAI
  • “Our [reinforcement learning]-trained attacker can steer an agent into executing sophisticated, long-horizon harmful workflows that unfold over tens (or even hundreds) of steps,” — OpenAI
  • “For most everyday use cases, agentic browsers don’t yet deliver enough value to justify their current risk profile,” — Rami McCarthy, Principal Security Researcher at Wiz