The Information Machine
Concluded·following since 5 Aug 2026·Day 1·3 sources

AISI finds Claude Mythos Preview and GPT-5.5 outperform all tracked autonomous cyber capability trends

The gist

Two frontier models have outperformed every trend line a government AI safety institute has measured for autonomous cyber task completion, raising questions about the pace of AI capability growth in this domain. AISI is also working with the UK's National Cyber Security Centre on how defenders can use frontier AI to stay ahead of attackers.

The full picture

The UK's AI Security Institute found that Claude Mythos Preview and GPT-5.5 have substantially exceeded the doubling trend in AI autonomous cybersecurity task performance AISI had been tracking since late 2024. AISI measures this via the '80% reliability cyber time horizon,' using the task duration a human expert would need as a proxy for AI autonomy. That metric's doubling time had already accelerated from roughly eight months in November 2025 to approximately five months earlier in 2026; both models have since outperformed any trend lines AISI has measured. In its evaluations, AISI found Claude Mythos Preview showed continued improvement on capture-the-flag challenges and significant improvement on multi-step cyber-attack simulations. GPT-5.5 ranked among the strongest models AISI has tested and is the second model to solve one of AISI's multi-step cyber-attack simulations end-to-end. Separately, Claude Mythos Preview identified a flaw in HAWK, a post-quantum cryptographic algorithm for identity verification, that halved its effective security strength, and sped up attacks on a deliberately weakened version of a common internet traffic cipher several hundredfold. Anthropic has released Claude Mythos Preview only to select organizations to identify and patch security flaws. Neither vulnerability directly endangers existing systems. Whether the capability jump represents an isolated event or the start of a faster trajectory remains unclear.

How it developed
5 August 2026

Claude Mythos Preview reported to have identified a flaw in HAWK post-quantum algorithm that halved its effective security strength

Sources
Semafor Technology
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free