Continuously hardening ChatGPT Atlas against prompt injection
OpenAI Blog
December 21, 2025 at 07:00 PM
📌 OpenAI is strengthening ChatGPT Atlas against prompt injection attacks using automated red teaming trained with reinforcement learning. This proactive discover-and-patch loop helps identify novel exploits early and harden the browser agent’s defenses as AI becomes more agentic.
OpenAI is strengthening ChatGPT Atlas against prompt injection attacks using automated red teaming trained with reinforcement learning. This proactive discover-and-patch loop helps identify novel exploits early and harden the browser agent’s defenses as AI becomes more agentic.
Read the full article at
OpenAI Blog
More in AI Tech
The AI industry has taken a doomer turn. What now?
This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development o...
AI agents blew the whistle on their cheating colleagues
A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, co...
AI Tech
DevFest is back
DevFest 2026 is back and here’s how you can connect with one of the more than 800 global events to build, secure, and scale in the agentic AI era.
AI Tech
MIT spinout turns plastic waste into resilient building materials
Atlas Building Composites is commercializing MIT research to turn plastic waste into parts for buildings and other infrastructure.