OpenAI built GPT-Red, an in-house AI hacker that attacks its own models to harden GPT-5.6 against prompt injection, and it works too well to release.

Exclusive: The firm says it wants to future-proof its safety procedures and stay ahead of human attackers.

OpenAI said its automated red-teaming model, GPT-Red, uncovered vulnerabilities that were used to make GPT-5.6 more resistant to attacks.