Public discussion of artificial intelligence has been dominated by digital applications: conversational agents, code generation, and image interpretation. However, a parallel transformation is occurring in the life sciences. Biology is shifting from a discipline that is primarily observed and described to one that is increasingly engineered. In this article, we review the evidence that machine learning has begun to treat the machinery of life as a design space, and we consider what that shift may mean for medicine and for computing.
I. Molecular Design: Rewriting the Instruction Set
Proteins constitute the functional machinery of the cell, and DNA encodes the instructions from which they are built. For most of the modern era this code could be read but not reliably rewritten. That constraint has weakened substantially. Generative models are now used to design proteins de novo at atomic resolution, and the resulting sequences can be expressed and validated experimentally rather than remaining computational artifacts.
Machine Learning and Gene Editing
A principal safety limitation of CRISPR-based editing is the occurrence of off-target effects, in which an editing enzyme modifies unintended genomic sites. Enzymes designed by protein language models have been proposed to reduce this risk. OpenCRISPR-1, an AI-designed nuclease, has been reported to produce approximately a 95% reduction in off-target editing compared with a conventional Cas9 reference. It should be noted that such comparisons are typically made in a limited number of cell lines and target loci, and generalization to clinical contexts remains to be established.
Development Timelines
Therapeutic development has conventionally required years to decades. In a widely reported case, a bespoke base-editing therapy for an infant with a rare urea cycle disorder was designed, manufactured, and administered within approximately six months. This represents a single-patient demonstration rather than a controlled trial, and the timeline achieved under compassionate-use conditions may not be reproducible at scale.
II. Economics of Discovery: From Screening to Design
Drug discovery has historically depended on high-throughput screening of large compound libraries, an approach with low prior probability of success per candidate. Machine learning methods that simulate molecular interaction and predict binding affinity are shifting part of this search from empirical screening toward computational design. Several quantitative trends have been reported in this context:
- Timeline compression: AI-assisted programs have advanced candidates from target identification to first-in-human trials in under 18 months, compared with a conventional interval of approximately 4 to 6 years.
- Phase 1 outcomes: Phase 1 success rates approaching 90% have been reported for AI-derived candidates, compared with a historical range of 40–65%. These estimates derive from small, non-randomized samples and are likely subject to selection bias.
- Program volume: The number of disclosed AI-designed drug programs has increased from three in 2016 to a projected 173 in 2026.
Capital allocation has followed a similar trajectory. In 2024, United States healthcare startups raised approximately $23 billion, of which close to 30% was directed to companies describing themselves as AI-native. It is worth noting that funding volume is an indicator of expectation rather than of validated clinical benefit, and the two have diverged in prior technology cycles.
III. Biological Computation
DNA is not only an analogy for code; it is an information-bearing polymer with measurable storage density. In terms of volumetric information density, biological media exceed any storage technology currently manufactured in silicon.
Information Density
The human genome comprises approximately 3 billion base pairs. When the theoretical limits of this medium are considered, the following estimates have been published:
- One gram of DNA can theoretically store on the order of 215 petabytes of data.
- At that density, the projected global digital archive of approximately 175 zettabytes would correspond to roughly 800 kilograms of DNA, occupying a volume of a few liters.
These figures describe theoretical capacity under idealized encoding assumptions. Practical read and write throughput, error rates, and synthesis cost currently remain orders of magnitude away from what would be required for general-purpose storage.
Energy and Analog Computation
The energy requirements of large digital models scale unfavorably with parameter count. The biological foundation model ESM3, with 98 billion parameters, required approximately 25 times more compute than its predecessor. By comparison, biological systems perform inference at markedly lower power; the human brain is estimated to operate on approximately 20 watts. This disparity has motivated interest in analog computation, in which the physics of the substrate performs the operation directly rather than being simulated by digital logic. Whether such architectures can be made programmable and reliable enough for practical use remains an open question.
IV. Data Volume in the Life Sciences
Sequencing cost has fallen by several orders of magnitude. The first human genome is estimated to have cost $3 billion to sequence, whereas current per-genome cost is approaching $100. As cost declines, data volume increases correspondingly: genomic research alone has been projected to generate between 2 and 40 exabytes over the coming decade. This scale makes automated analysis necessary rather than optional, since manual interpretation of billions of predicted structures and cellular interactions is not tractable.
"The first wave of applied machine learning automated the handling of information; the current wave is directed at the design of biological systems."
Conclusion
As molecular biology becomes increasingly programmable, therapeutic development acquires the character of a design problem rather than a search problem. The most consequential near-term effects are likely to be in target identification, protein engineering, and the compression of preclinical timelines, rather than in the wholesale replacement of experimental validation.
The convergence of machine learning and molecular biology is no longer speculative. It represents a change in how biological systems are described, designed, and modified.
Clinicians in the coming decade may use these models not only to interpret images, but to reason about interventions specified at the molecular level. The extent to which this capability translates into measurable patient benefit remains to be demonstrated.