OpenAI disclosed a self-replicating prompt injection that spread across simulated agent workflows, copying itself through outputs much like a computer worm.
The OpenAI Alignment report shows injected instructions propagating through email replies, filesystem artifacts, code comments, and a multi-hop Slack test. OpenAI said it observed no impact outside simulated tool calls in training and evaluation, so this is not a disclosed real-world incident.
The company is now adding self-reproduction to its attacker goals when training future models against prompt injection. For builders deploying AI agents, the practical risk is persistence: an instruction accepted in one step can survive in a file, message, or tool output and reappear later in the workflow.
