Skip to content

AI Models Begin Finding Cryptographic Breaks Beyond Published Research

Share
CryptanalysisBench benchmark graphic

Listen to this article

Read by Anchor

There is a difference between a model explaining a known attack on a cryptographic algorithm and finding the attack itself. The new CryptanalysisBench paper says that gap has begun to close rapidly, and that some models did more than reproduce the past.

In a test covering 191 problems across six families of cryptographic primitives, five advanced models broke most of the schemes already known to be vulnerable. They then moved into more difficult territory. They found practical weaknesses in schemes that had no known practical break within the test setting, producing results that the researchers believe may in some cases be new.

The benchmark is not looking for an elegant answer

Cryptanalysis is well suited to testing reasoning for a reason that is both useful and uncomfortable. A result is not judged by how persuasive it sounds or how closely the text resembles a reference answer. An attack either recovers a key, creates a collision or breaks a security property, and it can be run and verified.

The researchers divided the problems into three tiers. The first covered schemes with known practical breaks. The second covered schemes with no known practical break, tested both at full strength and in reduced versions. The third was a challenge set drawn from production cryptographic primitives at the boundary of published knowledge.

The five models were Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5 and the open-weights model GLM 5.2. In the first tier, they successfully broke between 65% and 86% of the schemes. In the second, they broke between 6 and 12 schemes at full strength, along with between 24 and 61 reduced versions.

Editorial illustration of a beam of light exposing a weakness in a cryptographic structure
Editorial illustration of a flaw found in a cryptographic structure · AI-generated image · AI Gossips

From reproducing results to finding a new flaw

The two leading examples in the paper are a key-recovery attack that exploits a design flaw in SpoC AEAD, and an error in the published security proof for the KINDI system against chosen-ciphertext attacks. The authors say that, to the best of their knowledge, neither result was previously known.

That scientific qualification matters. The paper was published on arXiv on July 20 and has not yet received enough public scrutiny for every claim of novelty to become an established fact. Finding a break in a candidate scheme or research design also does not mean that the cryptography protecting banks, phones and the web will collapse tomorrow.

Ignoring the result would also be a mistake. Modern cryptography depends on slow and expensive human review, and a flaw in a design or proof can sometimes remain for years before the right person sees it. If an agent can operate tools, revise its assumptions and test attacks that can be verified automatically, it adds a fast researcher to a process previously constrained by the small number of experts.

The capability is defensive and offensive at the same time

The benchmark can be used to test candidate schemes before adoption, and the defensive case is direct. Ask several models to attack a design in its early stages, collect the failed and successful paths, then have an expert review them. Reduced versions may also help reveal a pattern that does not emerge when testing starts immediately with full parameters.

The same tools also reduce the cost of looking for vulnerabilities in existing schemes. Not every primitive is widely deployed, and not every break can be turned into a real-world attack. Still, the direction of travel puts pressure on a long-standing and comfortable assumption: advanced analysis will always require a small team of specialists and a great deal of time.

The paper itself does not claim that models have reached the level of the best cryptanalysts in every respect. Performance varies, and success on a carefully designed benchmark is no substitute for open-ended work on a mature primitive. Some results may also depend on an abundance of compute, experimentation and feedback that an ordinary user cannot access.

What should change now?

For security teams, the message is not to replace cryptography experts with a model. The practical starting point is to bring agents into the early review stage, with an isolated environment, verification tools and a complete record of every attempt. A claim that can be executed should matter more than a confident explanation.

For anyone choosing algorithms for a product, the old rule still applies. Do not build your own cryptography, do not treat a single model result as a proven break, and do not adopt a new design before broad review. What has changed is that broad review may now include an automated adversary that never sleeps.

CryptanalysisBench does not say that the internet's secrets have been exposed. It makes a more precise point: the machine has entered a room whose mathematical key humans thought was still theirs alone.

Don't miss the next story

Subscribe for updates