Last May, we published a deep dive on how AI is breathing new life into biological research. And today, the field of biology is standing on the precipice of that transformation. In February 2025, Arc Institute unveiled Evo 2, an AI model trained on over 100K species' DNA, capable of identifying disease-causing mutations and generating entirely new genomes. This breakthrough represents a fundamental shift from merely deciphering life's code to actively rewriting it from scratch.
Evo 2 was trained on 9.3 trillion nucleotides extracted from more than 128K complete genomes across bacteria, archaea, phages, humans, plants, and various eukaryotes. In tests with the BRCA1 gene, it achieved over 90% accuracy in predicting which variants might cause breast cancer.
"Our development of Evo 1 and Evo 2 represents a key moment in the emerging field of generative biology, as the models have enabled machines to read, write, and think in the language of nucleotides," explains Patrick Hsu, Arc Institute Co-Founder and Core Investigator.
The journey to decode DNA began in 1869 when Friedrich Miescher first isolated what he called "nuclein" from cells. By 1953, Watson and Crick had revealed DNA's double helix structure, and by the 1970s, scientists were sequencing complete genes. These discoveries laid the foundation for genetic engineering and eventually the Human Genome Project, which fully mapped our genetic blueprint by 2003.
DNA analysis and computing has a rich history dating to the 1960s when Margaret Oakley Dayhoff pioneered using computers for biology research. This computational approach accelerated dramatically with next-generation sequencing technologies, allowing scientists to generate massive genomic datasets.
Each breakthrough in understanding genetics sparked new questions about function, regulation, and manipulation. The tension between unraveling nature's complexity and developing tools to modify it has driven progress, influenced by ethical considerations, computing advances, and funding priorities.
Evo 2 represents a quantum leap in technical capability. The model can process genetic sequences of up to 1 million nucleotides at once, enabling it to understand relationships between distant parts of a genome. This eight-fold increase over its predecessor was enabled by an architecture called StripedHyena 2, developed with assistance from OpenAI's Greg Brockman.
Evo2 offers immediate and tangible benefits for researchers and clinicians. The model excels at identifying disease-causing mutations in human genes with over 90% accuracy, potentially eliminating countless hours of expensive lab work while accelerating both diagnosis and drug development.
Researchers can use Evo2 to design gene therapies with cell-specific targeting, reducing side effects. As co-author Hani Goodarzi explains: "If you have a gene therapy that you want to turn on only in neurons to avoid side effects, or only in liver cells, you could design a genetic element that is only accessible in those specific cells."
The Arc Institute has prioritized accessibility by providing three user interfaces: a developer-friendly GitHub repository, integration with NVIDIA's BioNeMo framework to accelerate scientific research, and Evo Designer—a user-friendly interface requiring minimal technical expertise.
"We think of this as enabling an app store for biology," says Patrick Hsu, emphasizing that Evo2 functions as a foundation layer upon which specialized applications can be built. The fully open-source approach allows researchers to fine-tune the model for specific genes or organisms.
The significance extends beyond current capabilities. As Chief Technology Officer Dave Burke notes, researchers will likely discover "beneficial uses for Evo2 we haven't even imagined yet," from predicting how mutations affect protein function to designing genetic elements with novel properties. These innovations could revolutionize everything from disease treatment to sustainable bio-manufacturing.
As Evo models continue to scale, they may soon analyze a person's entire genome to predict disease risks, design targeted therapies, and perhaps craft biological structures never before seen in nature. The path from reading life's code to writing it has only just begun.



