CryptanalysisBench is a new benchmark designed to measure whether large language models can contribute to cryptanalysis, a field concerned with finding weaknesses in cryptographic systems. Developed by researchers affiliated with ETH Zurich, Anthropic, Tel Aviv University, and the University of Haifa, the benchmark provides a structured way to evaluate model performance across cryptographic problems with different levels of difficulty and real-world relevance.
The project matters because stronger AI systems could affect cryptography in two directions. They may help researchers identify weaknesses before schemes are deployed, but they could also change the practical risk profile of systems that have long been considered secure enough for ordinary use. The researchers position CryptanalysisBench as a way to track that capability shift rather than assume it has already occurred.
The CryptanalysisBench preprint on arXiv describes a three-tier evaluation framework, experimental prompts, and results for several frontier models. Its central contribution is not a claim that LLMs have broadly solved cryptanalysis. Instead, it creates a repeatable testbed for measuring progress on defined cryptanalytic tasks and for identifying where current systems do and do not perform well.








