---
title: "OpenAI's New Security AI Passes Advanced Hacking Tests 95% of the Time. Removing Safety Filters Only Got the Old One to 2%."
date: 2026-08-13T09:00:00Z
author: Deepankar Bhadrasen
url: https://truehorizon.ai/news/openai-gpt56-cyber-daybreak-security-model
description: "OpenAI's new specialized cybersecurity model found a Chrome zero-day. Benchmarks prove dedicated security AIs outperform general models with relaxed safety filters."
---

# OpenAI's New Security AI Passes Advanced Hacking Tests 95% of the Time. Removing Safety Filters Only Got the Old One to 2%.

> OpenAI's new specialized cybersecurity model found a Chrome zero-day. Benchmarks prove dedicated security AIs outperform general models with relaxed safety filters.

On August 10, OpenAI launched [GPT-5.6-Cyber](https://techcrunch.com/2026/08/10/as-ai-led-attacks-multiply-openai-launches-a-new-cyber-model/), a model purpose-trained for offensive security work: zero-day discovery, exploit development, vulnerability research. The headline number is that it completes 95% of tasks on OpenAI's own advanced cybersecurity benchmark. The more useful number is what happened when OpenAI tried the cheaper alternative first: just removing the safety filters from its general-purpose model instead of training a new one.

## What GPT-5.6-Cyber actually is

GPT-5.6-Cyber sits behind a new tier OpenAI calls Daybreak Red: tightly vetted access limited to trusted partners including Accenture, IBM, CrowdStrike, and Cloudflare, gated by identity verification, legal attestations, monitoring, and hardware security keys for individual accounts starting September 1. A lighter tier, Daybreak Blue, gives approved defenders access to the general-purpose GPT-5.6 Sol model for everyday work like incident response and malware analysis, with some safety guardrails relaxed for verified users.

## The gap that actually matters

Here's the comparison worth sitting with. Standard GPT-5.6 Sol, with its normal safeguards, completes 1.5% of tasks on OpenAI's advanced cybersecurity benchmark. Unlocked under Daybreak Blue, with some of those guardrails removed for vetted defenders, it reaches 2.0%. Barely moved. GPT-5.6-Cyber, purpose-trained on the same category of task, completes 95.0%. Removing a safety filter and training a specialized model are not the same intervention, and OpenAI's own numbers make the difference impossible to miss.

## Proof in the wild

This isn't just a benchmark claim. GPT-5.6-Cyber found CVE-2026-15903, a real, previously unknown compiler flaw in Chrome's V8 engine that let memory be read and overwritten inside the browser's sandbox. Google patched it through coordinated disclosure. OpenAI says additional undisclosed vulnerabilities in mobile operating systems, databases, and OS kernels are moving through the same responsible-disclosure process.

## The part OpenAI felt it had to clarify

Alongside the launch, OpenAI stated directly that "GPT-5.6-Cyber was not involved in the exploitation of Hugging Face, and no model with that involvement is slated for release." That's a notable thing to feel compelled to say. A model this capable at offensive security is exactly the kind of tool people should ask hard questions about, and the fact that OpenAI addressed it head-on suggests they know it too. The identity checks, legal attestations, and hardware security key requirement around Daybreak Red aren't compliance theater here. They're the actual containment mechanism, doing real work rather than sitting on a slide.

## What this means if you're evaluating AI security tools

Three things worth doing before you take a vendor's AI security capability claim at face value:

- Ask for the purpose-trained benchmark next to the general model's benchmark on the same tasks. If a vendor can't show that gap, they probably haven't measured it.
- Treat access controls (vetting, identity checks, monitoring) as part of the product you're evaluating, not paperwork bolted on afterward. Weak gating around a highly capable model is its own risk.
- Assume the same capability curve applies to attackers. A model that can find zero-days this reliably shrinks the window between a vulnerability existing and someone exploiting it, on both sides of that equation.

## Where TrueHorizon fits

This is exactly the distinction we help clients make: what an AI tool actually does under real constraints, versus what a demo or a marketing claim suggests it does. **That's not a checklist we're learning on your project.** It's the expertise we bring to it. Whether it's a security model, an agent framework, or a vendor's usage metric, the question is the same: does the capability hold up once you look past the number they led with.

If you're evaluating AI tools for a security or compliance-sensitive function, [take our AI readiness assessment](https://truehorizon.ai/assessment) before you take a vendor's capability claim at face value.

Source: https://truehorizon.ai/news/openai-gpt56-cyber-daybreak-security-model
