But most of the non-coding DNA is junk—the remains of ancient viral infections, DNA-level parasites, genes that have been inactivated by mutation, and so on. Figuring out what’s useful and what’s not has been an ongoing challenge for biologists for many reasons.
First, the proteins that interact with DNA aren’t that picky about the sequences they stick to, potentially binding at random throughout the genome and tolerating a certain degree of mutation. Many of these proteins are also cell-type specific; there’s a different population of DNA-binding proteins in liver cells, nerve cells, immune cells, and so on. In many cases, having many different protein binding sites in a compact space matters more than the presence of any one of them.
We’ve developed various software tools that identify individual sites of interest in non-coding DNA. But this is exactly the sort of problem that AI is good at solving: one involving probabilities that are imprecise and rely heavily on context. So Google developed the AlphaGenome AI system, which evaluates sequences for their potential function.
(For the biology geeks that don’t want to sort through the paper, AlphaGenome attempts to identify “gene expression, transcription initiation, chromatin accessibility, histone modifications, transcription factor binding, chromatin contact maps, splice site usage, and splice junction coordinates and strength.”)










