Red to hack its own AI, and hid it
OpenAI developed GPT-Red, an AI designed to find vulnerabilities in other AI systems. This automated red-teamer uses self-play to discover new attack methods, like 'fake chain of thought,' and has demonstrated high success rates against older models. OpenAI is keeping GPT-Red private to prevent misuse, acknowledging that human expertise remains crucial for AI security.