Security Is a Big World

On September 1, Anthropic shipped Claude Fable 5.1 and Mythos 5.1: same underlying model, different safeguards. Fable is broadly available. Mythos is gated to organizations vetted through Anthropic’s Cyber Verification Program. Two days later, OpenAI shipped Astra, its first model rated Critical for cyber. Its strongest capabilities are likewise restricted to vetted defenders via Daybreak.

Two frontier labs, two days apart, one structure: ship defensive capability broadly, gate its highest-risk forms. I expect these gates to widen. Access broadens anyway: open weights trail frontier cyber capability by four to seven months, per UK AISI.

The bottleneck was never discovery

Cyber disclosure is a coordination problem. A study of 8,073 open-source projects published in ACM Transactions on Software Engineering and Methodology found more than 80% of practitioners supported coordinated disclosure. Yet only 55% of vulnerabilities followed it, a median 36 days from patch to disclosure. That is open-source practice, not vendor disclosure. But the failure mode is familiar: the fix can be public before anyone assessing exposure knows what it fixes.

The problem gets harder for AI systems

Longpre and colleagues argue AI flaw-reporting norms lag software security. Their FLARE-AI follow-up names the problem: whoever discovers a flaw may need to report it separately to every affected stakeholder.

And increasingly capable models keep shipping into that world.

Security is a big world

The Big World Hypothesis offers the lens: for many learning problems, the world is far larger than any agent. More capability does not close that gap. Dependencies multiply; no defender or model sees the whole system.

Security is a big world. Every defender in it is small, and so is every model.

Lasting advantage will not come from the agent that sees the most, but from moving partial knowledge faster between people and agents that see different slices.

That is the harness: workflows, disclosure channels, verification steps, and human judgment that turn intelligence into action. It does not benchmark cleanly. But I think the harness is where defensive advantage lives.

Announcements

Sources

  • Longpre et al., “FLARE-AI: Flaw Reporting for AI,” arXiv:2606.31567 link
  • “An Empirical Study on Vulnerability Disclosure Management of Open Source Software Systems,” ACM Digital Library link
  • Javed and Sutton, “The Big World Hypothesis and its Ramifications for Artificial Intelligence” link
  • UK AI Security Institute, “How Far Behind the Frontier are Leading Open Weight Models on Cyber?” link

Similar Posts