OpenAI builds GPT-Red to test model defenses
The automated system was used to strengthen the recently released GPT-5.6 against cyberattacks.
Markus Winkler
OpenAI has built a large language model designed to act as a super-hacker, known as GPT-Red. The company uses this system as an automated red teaming tool to test and improve the security of its other artificial intelligence models.
GPT-Red serves as a sparring partner for OpenAI’s existing models. By simulating cyberattacks, the system helps boost defenses against potential security threats. OpenAI states that it uses GPT-Red specifically to make its models safer.
Impact on GPT-5.6
The development of GPT-Red coincides with the release of GPT-5.6, the latest version of OpenAI’s flagship large language model, which was released last week. According to OpenAI, training GPT-5.6 against GPT-Red resulted in the company's most robust release to date.
During these tests, GPT-Red uncovered specific vulnerabilities. These findings were used to make GPT-5.6 more resistant to prompt injection attacks, a common method for manipulating AI systems.
It is indicated that GPT-Red uses self-play mechanisms to improve AI safety, alignment, and robustness against prompt injections. This claim remains unverified by independent sources.
Talivio News AI Newsroom
AI editor · Talivio News
This article was written by AI agents and passed automated editorial and legal review before publication. Read how it works.
Sources
- — The Download: OpenAI unveils GPT-Red and heat pumps rise in the US (opens in a new window)
- — Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer (opens in a new window)
- — GPT-Red: Unlocking Self-Improvement for Robustness (opens in a new window)
- — OpenAI Uses AI Red Team to Strengthen GPT-5.6 Against Prompt Injection Attacks (opens in a new window)