As cited
Copy frozen at (site build).
research
Measuring LLMs’ Ability to Perform Cryptanalysis
Anthropic and collaborators released CryptanalysisBench, a benchmark of 191 cryptanalytic tasks spanning historical and modern cryptographic primitives, to measure large language models' ability to discover mathematical attacks. Five frontier models successfully broke 65-86% of known vulnerable schemes and discovered novel attacks, including a key-recovery vulnerability in the SpoC AEAD cipher and an error in KINDI's security proof, neither previously known.
Why it matters: Security practitioners and cryptography teams need to monitor AI progress in cryptanalysis as a potential long-term threat to cryptographic systems currently used in production, particularly as models demonstrate capability to find previously unknown vulnerabilities in standardized algorithms.
- Source published
- First seen by Cybersecurity Tracker
Source attribution
Correction
Correction recorded as of .
research
Measuring LLMs’ Ability to Perform Cryptanalysis
Anthropic and collaborators released CryptanalysisBench, a benchmark of 191 cryptanalysis tasks to measure large language models' ability to discover mathematical attacks against cryptographic algorithms. Frontier models including Claude Opus 4.8 succeeded in breaking known vulnerable schemes and, notably, discovered new attacks including a key-recovery flaw in SpoC AEAD and an error in KINDI's security proof. The benchmark serves as both a reasoning testbed and an early warning system to identify weaknesses in cryptographic primitives before production deployment.
Why it matters: Cryptography practitioners and standards bodies should monitor whether LLMs develop practical cryptanalysis capabilities that rival or exceed human expert analysis, as this could accelerate discovery of vulnerabilities in deployed and candidate algorithms.
- Source published
- First seen by Cybersecurity Tracker