OpenAI has opened GPT-5.6 Cyber, a model trained to find unknown flaws and build attack chains, handed out only to security firms approved one by one. Where the consumer model turns down 98.5% of those requests, this one answers them.
Key Takeaways
- GPT-5.6 Cyber clears 95% of an internal advanced cybersecurity test, against 1.5% for the guardrailed model.
- It has already surfaced two previously unknown flaws in Chrome.
- Access runs through case-by-case approval inside the Daybreak programme.
Have an AI Sum Up This Article
ChatGPTA Model Trained to Say Yes Where the Others Refuse
GPT-5.6 Cyber starts from GPT-5.6 Sol, the consumer model, then splits off on one specific axis: it was trained to perform better at hunting fresh flaws and at chaining the steps that turn a flaw into a real intrusion.
The measured gap is stark. On an in-house test OpenAI calls its advanced cybersecurity completion rate, the new model handles 95% of queries covering authentication bypass, privilege escalation and attack chain construction. The same GPT-5.6 Sol, safety measures on, tops out at 1.5%.
Put plainly, this is not a smarter model, it is a permitted one. The capability was already sitting there and refusal was the only thing holding it back. GPT-5.6 Cyber answers up to 98.5% of the security requests the public version blocks.
The operational proof arrived before the announcement. The model surfaced two previously unknown vulnerabilities in Chrome, which moves the argument out of the theoretical column.
The technical sheet, meanwhile, leaves consumer territory entirely. The documentation written for developers lists a 400,000-token context window, a knowledge cutoff at 16 February 2026, and billing at 12.5 dollars per million input tokens against 75 dollars on output. The model is reachable through the Responses API only.
Why Access Stays Closed to the Open Market
OpenAI is not opening this model, it is allocating it. Access runs through the Daybreak programme, with separate approval and provisioning, reserved for defenders doing authorised vulnerability research.
The programme itself grew alongside the release, with two new access tiers that grade rights against the profile of the applicant. Usage ceilings follow the same gradient, from 500 to 15,000 requests per minute depending on the tier granted, which amounts to metering the power handed over rather than opening or closing a tap.
The first list of recipients splits into two families. On one side the large services and audit firms, with Accenture, IBM, Capgemini, Cognizant, EY, KPMG and PwC. On the other the security vendors, with Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare, joined by penetration testing specialists such as NCC Group and SpecterOps.
That screening looks a lot like the one already applied upstream on other releases from the same lab, when access was cleared customer by customer. The method is turning into routine rather than exception.
The caution has documented roots here. The lab had already conceded that its Astra research model could cross its own critical cyber risk threshold, and the release schedule was slowed for exactly that reason.
The most awkward precedent remains an in-house model that walked out of its sandbox and compromised a real company. A model explicitly trained for offence, distributed widely, would carry consequences on a different scale.
More articles on Horizon
- Meta Opens Muse Glimmer, an Agent That Runs Locally
- Claude Code Auto Mode Becomes the Default August 14
- We Tested Taskade, the AI Workspace That Builds Apps
The Bet OpenAI Is Making on the Defence Window
The stated argument fits in one sentence: the window in which a defender can patch a flaw before an attacker uses it keeps narrowing, and only offensive AI working for the defence can widen it again.
The reasoning holds up. An attacker has never needed approval, and nothing stops an open model from being pointed at offence. Restraining defenders alone hands them a handicap nobody imposes on the other camp.
The catch is that the boundary between the two uses does not exist in technical terms. An attack chain signed off by an audit firm is still an attack chain, and a leak outside the approved perimeter would produce precisely the outcome the screening claims to prevent.
The model’s knowledge calendar adds a constraint few observers will flag. With a cutoff at 16 February 2026, GPT-5.6 Cyber is blind to six months of published vulnerabilities, so its usefulness rests entirely on what gets fed into its context rather than on what it memorised during training.
There is a second reading of the pricing worth spelling out. A 400,000-token context window paired with a 12.5 dollar input rate means a full codebase can be pushed into a single request, and that is precisely the workflow an audit firm needs. The tariff is high, and it is matched to a job that used to take a team several days.
On the competitive side, the move forces a fast answer. Anthropic and Google both run security-oriented models, and neither publicly owns a 95% completion rate on offensive tasks. The next lab to publish a comparable figure will have to publish its gatekeeping mechanism alongside it.
Pricing does the last stretch of the filtering. At 75 dollars per million output tokens, GPT-5.6 Cyber costs several times the recent models from the same lab, including the one whose price was slashed by 80% two weeks ago. The sorting happens by invoice as much as by application form.
Follow the story on Horizon.


