The newly released AlphaGenome Atlas tries to fill that gap with predictions. For each of those nine billion changes, it offers an estimate of how the change would likely affect molecular processes across hundreds of cell types and tissues. The dataset spans one petabyte, making it more than 30 times the size of the AlphaFold database for protein structures.

It builds on the AI model AlphaGenome, introduced in 2025. The model reads DNA stretches one million letters long and predicts how strongly a gene gets read, whether regulatory proteins can bind to the DNA, and how a gene's transcript gets spliced. Until now, the model had to be queried for each variant one at a time. Now the answers are precomputed. According to the paper, each variant comes with about 27,000 individual prediction values on average.

That matters most for the roughly 98 percent of the genome that holds no blueprints for proteins. These noncoding regions act like switches and dials that decide when and in which tissue a gene is active. That's where most disease-linked variants sit, and it's also where their effects have been hardest to read.

One number for every mutation

Thousands of prediction values per variant are too much for everyday use, the team says. So Deepmind built the AlphaGenome Variant Impact Score (AVI), which boils it all down to a single number. A small neural network combines the AlphaGenome predictions with the protein model AlphaMissense and two measures of how unchanged a DNA site has stayed across millions of years of evolution. AVI works with 18 input features. The established benchmark tool CADD uses more than 150.