OpenAI built a model that hacks like a pro. Even OpenAI is nervous
Here's a product launch you don't see often: a company ships its new flagship AI while publicly warning everyone what it's capable of.
That's how Astra — the model behind GPT-6 — arrived. It's the first system to cross OpenAI's internal "Critical" threshold for cybersecurity capability, the top of the company's risk scale, a tier no earlier model touched. Translated from framework-speak: Astra can find security holes nobody knew existed, then exploit them, without a human guiding it step by step.
What happened in testing
The lab results explain the label. During internal evaluations, Astra dug up two zero-days — flaws unknown to anyone — and chained them into a complete browser-compromise attack that broke out of its sandbox and ran commands on the host machine. In another test it built a privilege-escalation chain and reached root on a hardened operating system.
The chaining is the scary part. AI finding individual bugs? Old news. Stringing fresh discoveries into a working, multi-stage attack has always been what separated elite human hackers from script kiddies. Astra does the elite version, on demand.
Sell the model, cage the feature
OpenAI's answer is a careful straddle. Astra itself is rolling out widely. Its most dangerous skills aren't. The advanced offensive-security capabilities are limited to a handful of vetted organizations inside OpenAI's Daybreak enterprise coalition — a small alpha group first, wider vetted access later, general public never (for now).
OpenAI researcher Fouad Matin summed up the bind without spin: these capabilities "can and will help defenders find and fix serious weaknesses," but without safeguards "they could also make attackers more effective." Both halves of that sentence are true, which is exactly the problem.
Credit where due: OpenAI spent August tightening its cyber controls before any of this shipped, so the gating wasn't a scramble. But stop and notice the moment — a tech company fencing off its own headline feature at launch. That's not how this industry usually behaves.
The in-house doubts got louder, too
The nerves go beyond one gated feature. Days after launch, OpenAI's chief scientist Jakub Pachocki published a warning of his own: no lab has solved alignment and monitoring well enough to keep scaling at full speed, and he's concerned "no one is prepared" for a faster loop of AI building AI.
Read that again with the org chart in mind. The person running the science says the safety tooling is behind the capabilities. At that point "Critical" isn't a marketing tier. It's a confession with a version number.
What it means for you
If you run security anywhere, this cuts both ways. The exact skill that spooked OpenAI's evaluators could let defenders attack their own systems the way real adversaries would — creatively, relentlessly, at machine speed. The catch is access: vetted Daybreak members get that power first, and everyone else gets to hope the fence holds.
For regular people, the stakes are quieter but real. Your browser, your OS, every service you log into — all of it now lives in a world where finding vulnerabilities is becoming an industrial process. If the gating works, the good guys patch faster than the bad guys exploit. If it leaks, the best hacking tool ever built comes with a login page.
Nobody knows which way it breaks yet. What's settled is the milestone itself: an AI crossed the line from helping hackers to being one. We know because its maker said so out loud.
Image: Tima Miroshnichenko, via Pexels





