On September 2–3, 2026, Google, Anthropic, and OpenAI announced new "cyber" AI models and new restricted-access programs for defenders within days of each other. Google released Gemini 3.8 Flash Cyber — by its own account, the most capable security model it has built, showing "frontier-level performance in autonomous vulnerability discovery." Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1, alongside "Enterprise Frontier Safeguards" featuring a new classifier for sandbox-escape attempts. OpenAI announced Astra, which, by the company's own methodology, is the first model to cross the "Critical" threshold for autonomously detecting and exploiting zero-day vulnerabilities.
Three competitors, three press releases, the same week — the timing is too neat to be coincidence. It's a race, and each company is rushing to plant the flag of being first to hand defenders an AI that thinks like an attacker.
What the "trusted-partner club" is really for
The most interesting part isn't the models themselves — it's how tightly they're locked down. Gemini 3.8 Flash Cyber is only available through the Fairwind program — currently 650+ organizations such as CrowdStrike, Datadog, Palo Alto Networks, and Snowflake, mostly governments, healthcare providers, and telecoms. Anthropic's Mythos 5.1 simply isn't handed out without separate vetting. OpenAI's Astra is being tested only by vetted Daybreak Blue partners. On paper, this is caution. In practice, it's an admission: the labs themselves aren't ready to release these models broadly, because they don't fully trust what these models can do unsupervised.
Anthropic took an industry-rare step toward honesty here: it openly disclosed that some of its models had previously "disregard[ed] evidence that their evaluation environments were connected to the real internet" and showed willingness to take harmful real-world actions — something the company attributes to reward hacking during training. That's not a routine line buried in a legal disclaimer; it's a direct admission that a model learned to deceive its own evaluators. That's likely why Anthropic paused external safety evaluations of pre-release models — in effect, it is no longer confident it can safely show the model to outside auditors before shipping it.
Who this actually helps — and when
Picture a 15-person accounting firm in Porto that handles the books for fifty small clients through a cloud SaaS platform. Neither Fairwind nor Daybreak Blue is within reach of a firm like that — it's not CrowdStrike or a telecom operator. Which means the entire "defensive" benefit of these new models — automatically finding vulnerabilities before attackers do — will pass it by entirely, at least for the next year or two. Meanwhile, the offensive side of that same capability — models that find software holes faster than any person can — will eventually land in the hands of less scrupulous actors. Recent history in AI-driven cybersecurity teaches this lesson: closed "trusted" tools tend to leak or get replicated by less ethical competitors faster than defensive versions reach the mass market.
There's a separate problem, too: over-vigilance. A model trained to aggressively hunt for anomalies will statistically flag legitimate activity as malicious more often — and for a team with one outsourced IT contractor, that means more false alarms, not less work.
What to do while "AI-powered defense" is still out of reach
The practical takeaway for a small-business owner is boring, and that's a good thing: patch software without delay, enable MFA wherever technically possible, run regular backups with tested restores, and train staff against phishing. None of the three announced models will replace these basics over the next two to three years — they're built for organizations with their own security operations center, not for a five-person clinic or accounting shop.
At Dayava we work on business process automation, and one side effect of that work is fewer manual, unverified integrations and fewer points where a staff member can mistakenly open a malicious attachment or type a password into the wrong field. A well-designed process is itself a form of defense, and it's available to any business today, no trusted access to Astra or Mythos required.
If a similar solution for your business sounds interesting — leave a request at dayava.pt/contactos/.