On Monday, Sept. 28, 2026, Anthropic released Claude Sonnet 5.5, a faster and cheaper model built for everyday business work.
Anthropic says it does not advance the frontier of its capabilities. Yet it ships with cyber safeguards the company once kept for its most powerful systems.
The reason is blunt. Anthropic said the model’s cyber skills are a “large improvement” over its predecessor and now rival Claude Opus 5, a more expensive model. Higher-risk requests, which could include penetration testing and exploit generation, now get handed to the older Sonnet 5, a company representative told The New Stack.
Cyber skills are reaching cheaper models within weeks
Sonnet 5.5 delivers comparable cyber skill at token prices 60% lower, based on Anthropic’s published rates, and Opus 5 only debuted on Friday, July 24, 2026. A capability the industry treats as dangerous is now getting cheaper on a timeline measured in weeks.
Anthropic did not plan it that way. The company said it deliberately avoided training Opus 5 on cyber tasks, yet the model improved anyway as it grew smarter overall. If that pattern holds, no lab can build a better coding assistant without also sharpening its offensive potential. Artificial intelligence is now at the stage of self-improvement.
The safeguards are climbing just as fast. Sonnet 5 already carried lighter cyber controls modeled on older Opus versions, Yahoo Finance reported in June. Sonnet 5.5 now gets flagship protections.
That complicates the debate CEO Dario Amodei started. In an essay published on Saturday, Sept. 12, 2026, he wrote that “we must slow the pace at which we improve the capabilities of AI models.”
Sonnet 5.5 is the company’s second launch since Dario’s “slow down” essay. The pace that matters for security may be how fast top capability trickles down, not how fast the top rises.
Anthropic says Claude Sonnet 5.5 has cyber skills comparable to Opus 5, making it the first Sonnet to launch with safeguards built for its top models.Leon Neal / Getty Images
Guardrails that stop attackers can also stall defenders
While top-tier safeguards address Amodei’s concerns, these safety filters carry a cost that rarely makes a launch post.
In July, Hugging Face said guardrails on a leading U.S. model blocked its security team from analyzing an attack by autonomous AI agents, Fortune reported. The team switched to GLM 5.2, an open model from China’s Z.ai, running on its own servers.
Hugging Face’s incident recap, as quoted by Constellation Research, framed the imbalance plainly: the attacker “was bound by no usage policy.” Defenders relying on hosted models had no such freedom.
Related: OpenAI paused training again, and Washington’s in the middle of it
That friction now reaches Anthropic’s midtier model.
“Sonnet is really for the cost-conscious customer,” Theo Chu, a research product manager at Anthropic, told CNBC.
API developers must opt in to automatic fallback, according to The New Stack, so companies building on Sonnet 5.5 must plan for that routing.
Sonnet 5.5 is strongest at well-scoped work like spreadsheets, slide decks and polished documents, according to Anthropic. It even beat Opus 5.5 on Terminal-Bench 4.0, an agentic coding test, in company testing.
It falls short on complex, open-ended work needing sustained judgment, where Anthropic says Opus 5.5 stays clearly stronger. Some microbiology and virology requests may also get flagged in error, the company warned.
Anthropic says routine bug fixing stays open. Verified defenders can soon apply for tiered access through an expanded Cyber Verification Program.
Verified access could favor the biggest security firms
This shift toward verified access directly affects the market.
Cybersecurity stocks have traded on Anthropic headlines all year. CrowdStrike (CRWD), Palo Alto Networks (PANW) and Zscaler (ZS) each fell more than 5% on Friday, March 27, 2026, after a report raised concern that Anthropic’s Mythos model could help hackers skirt current defenses, according to Bloomberg. The market’s first instinct treated stronger AI as a threat to incumbents.
That instinct has since flipped. CrowdStrike CEO George Kurtz later credited the Mythos scare for the best quarter in company history.
More Anthropic:
Anthropic just handed Akamai a game-changing deal
Google, OpenAI, and Anthropic just made a move on AI safety
Anthropic researcher resigns and his reason is a warning to us all
Cybersecurity shares then rallied on Monday, Sept. 14, 2026, even as AI stocks sank on Amodei’s slowdown call, CNBC reported.
Verified access pushes that shift further. CrowdStrike and Palo Alto already help test unreleased models from Anthropic and OpenAI, CNBC reported.
As the strongest cyber tools move behind identity checks, vendors inside that circle gain a head start that smaller rivals and in-house teams may struggle to match.
Safety is shifting from the model to the customer
Anthropic is no longer limiting what its models can do. It is now deciding who gets to use its sharpest cyber skills. OpenAI already splits its own cyber program into two access tiers, each requiring identity verification.
The result looks less like a software market and more like banking, where lenders must verify customers before opening accounts. The strongest capabilities go to approved users first.
That matters for investors because Anthropic has no public shares yet. The company confidentially filed for an IPO on Monday, June 1, 2026, and public investors will eventually value a business that sells its most sensitive capabilities through verification, not just a price list.
That framework could lock in enterprise security customers, or it could cap how fast those products grow.
Watch Haiku 5.5, the cheapest tier, which Anthropic says arrives in the coming weeks. If it also needs cyber gates, the word “frontier” will no longer tell regulators or investors where AI risk actually sits.
Related: Anthropic’s IPO is a $300 billion test for Amazon
