Stanford’s AI Designed Bacteriophages: Ushering in the Next Phase in IP Strategy for Phage Companies?

Craig Thomson
Partner at 

Craig Thomson is an Irish, UK and European Patent Attorney and Partner at HGF, where he leads the firm's multidisciplinary Microbiome IP Team. With extensive experience advising companies, investors and research organisations across the microbiome, live biotherapeutic, probiotic, diagnostic and nutrition sectors, he provides strategic intellectual property guidance from early-stage innovation through to commercialisation and investment. Craig has held a number of leadership and advisory roles within the microbiome sector and is recognised for his deep understanding of both the science and commercial dynamics shaping the industry

Roxna Kapadia
Patent Director at 

Roxna is a Patent Director and member of HGF’s Bioinformatics and Microbiome IP Teams. With a background in microbiology and immunology, she specialises in protecting innovations at the interface of life sciences and computational technologies, including bioinformatics, AI and machine learning. Having previously worked in-house within the pharmaceutical industry, Roxna brings a commercial perspective to patent strategy, alongside experience in microbiome technologies, gene editing, viral therapeutics and personalised medicine

Stanford’s AI-Designed Phages: Why This Matters

Researchers at Stanford University and the Arc Institute used Evo 2 (a generative AI model for DNA design) to create complete bacteriophage genomes. Starting from ΦX174, they synthesised nearly 300 AI‑designed genomes and found that 16 produced viable, bacteria‑killing phages. This is an interesting experimental proof of concept, but it is nowhere near clinical validation .

This point was echoed by Professor Martha Clokie, Professor of Microbiology and Director of the Becky Mayer Centre for Phage Research, who commented to us that:

This proof of principle study is important as it shows that genome context can be learned, but we need to keep biology in perspective. Even extensively evolved ΦX174 phages are poor therapeutic starting points because of their tiny, constrained genome and narrow biology that can’t accommodate the complex bacteria-recognition, anti-defence and infection machinery needed to address diverse, clinically important bacterial pathogens. The ‘rules’ learned from this recent study can’t be directly applied to designing the much larger and more complex phages that we would typically consider for therapeutic use.

It is however a notable upstream shift: instead of using AI to analyse or optimise natural phages. Stanford’s work does show that generative models can produce complete genome sequences which, in a small proportion of cases, yield viable phages. Such approaches may ultimately complement and extend phage discovery by learning from it, and helping us navigate the enormous diversity that evolution has already generated.

Current use of AI and bioinformatics technologies

AI and bioinformatics technologies are already embedded in therapeutic phage development. Locus Biosciences uses predictive AI, synthetic biology and proprietary genomic and clinical datasets to design engineered phage cocktails. SNIPR Biome applies genomic data and bioinformatics to identify bacterial fingerprints and programme CRISPR systems. Armata Pharmaceuticals relies on sequencing databases, software and comparative genomics to analyse and engineer natural phages. Phage product design is already data‑driven. Importantly, the usefulness of these approaches depends not only on computational capability but on the quantity and quality of biological data available to train, test and experimentally validate predictions.

The scientific advance in this field is the transition to AI designing functioning genomes. The corresponding IP question is whether protection strategies should follow the algorithm, the design process, the generated genome, the engineered organism or all four.

Conventional Protection strategy

Readers of our earlier article, “ELIGO v SNIPR: a reflection on IP strategies in a competitive environment” (Microbiome Times, 7 February 2022), will already be aware  that patent portfolios in this therapeutic field are plentiful and already subject to robust court battles. This first wave of patent filings has been largely directed to protecting the core biological innovations of this new class of therapeutic; tending to focus on the physical products of the technology (i.e. engineered phages, CRISPR payloads, defined cocktails and therapeutic uses).

Protecting the Platform, the Product and the Pipeline

We are now starting to see a second wave of patent filings, with increasing emphasis on AI- and bioinformatics-driven technologies. This second wave is part of an evolution in the IP strategies employed by this field. AI-driven platforms create an additional layer of potentially protectable subject matter in the computational workflow itself. An AI or machine-learning platform can be thought of as the “design engine” that links genome sequence to biological function. In practice, its value depends on having access to high-quality proprietary training data and experimental systems capable of validating predicted phenotypes. Protecting that design engine, rather than just the individual phages it produces, has the potential to provide broader commercial protection because it can underpin the ability to develop an entire pipeline of future therapeutics. For developers of generative genome technologies, the strongest IP position may therefore combine protection for the computational platform with more conventional patent claims to the resulting engineered biological products and their uses. Conventional patent claims may help prevent competitors from copying individual phage therapeutics or defined cocktails. By contrast, protection strategies directed to the computational platform may provide protection further upstream by depriving competitors of access to the efficient design approach needed to generate rival therapeutic candidates at scale. That raises a strategic question: what might Stanford and the Arc Institute seek to protect from this recent work?

Patent applications related to the Stanford work are unlikely to be visible yet, because patent applications are generally published 18 months after their initial filing date. It may be relevant to consider another patent filing that has been made previously by one of the inventors of the Stanford work, Brian Hie.   US Patent No.11011253 B1concerns “escape profiling for therapeutic and vaccine development”. The patent uses a language-model approach to analyse protein sequences and identify mutations that may permit viral immune escape. Importantly from a patent-strategy perspective, the claims do not stop at the abstract use of machine learning. They link training and application of the model to the generation of an escape profile, identification of biologically relevant regions and downstream vaccine or therapeutic design. This illustrates an important approach for AI-biology inventions: anchoring the computational analysis to a defined biological problem and a practical technical output, rather than attempting to claim an algorithm in isolation.

Four Layers of Potential Protection

Going forward, the potentially protectable subject matter in the AI-driven world will be broad and may include at least four layers of protection. First, there is the platform itself: for example, methods using a genome-scale generative model to generate candidate viral genomes subject to one or more biological constraints, followed by computational selection, synthesis and/or experimental validation. Ideally, such claims would not be tied too closely to a particular model, but would seek to capture the more general workflow by which whole genomes are generated to achieve a desired phenotype. Second, the resulting biological products may themselves support more conventional patent protection, including synthetic bacteriophages, nucleic acids encoding their genomes, sequence-defined variants, altered host-range determinants and other engineered components. Third, there may be product- and indication-specific opportunities around particular phages or combinations of phages that exhibit clinically useful properties, such as an expanded host range or the ability to overcome bacterial resistance. Finally, some of the most valuable aspects of the platform, including training-data curation, candidate-ranking methods, optimisation criteria, negative experimental data and iterative wet-lab feedback, may be better retained as trade secrets where they are difficult for competitors to detect or reverse engineer. However, each developer’s position will be different. The precise content of the IP strategy, and the appropriate balance between patent protection and trade-secret protection, will therefore need to be assessed carefully; this emerging as one of the most interesting new skills of patent attorneys specialising in this technical area.

Commercial takeaway

The commercial lesson is therefore not simply that AI may accelerate the identification, optimisation or generation of phage candidate but that this also opens up the opportunity to create a more layered IP estate around an AI-enabled biological design platform. In this emerging landscape, access to unique phage diversity and experimentally validated phage-host data may prove just as strategically important as ownership of the generative model itself.  Established companies such as Eligo Bioscience, SNIPR Biome, Locus Biosciences and Armata Pharmaceuticals illustrate the value of building protection around the tangible outputs of phage technology . Generative genome platforms potentially add another defensible layer upstream, covering the computational process by which those products are designed. The strongest strategy is therefore likely to combine platform claims, product claims and therapeutic-use claims, supported by selective use of trade-secret protection.

This point was emphasised by Professor Martha Clokie who commented to us that:

I think the really exciting opportunities are for AI to learn from the extraordinary diversity that nature has already evolved and see how that could be exploited. So large, well-characterised phage collections linked to detailed biological data could become enormously valuable, both scientifically and commercially, as they are needed to give the training and validation data needed to make AI-designed phages useful.

Biosecurity: The Next Layer of Protection

There is, however, one further layer to consider: biosecurity. As this field develops, credible biosecurity technologies and protocols will need to become an integral part of the technology stack. Training-data governance, sequence screening, controlled synthesis, auditable safety filters and related safeguards will not be peripheral compliance measures; they may become enabling technologies without which AI-designed phage products cannot be responsibly developed, manufactured or accepted in the market. That creates a further IP opportunity. If a particular biosecurity technology were to become the trusted or industry-standard safeguard for AI-enabled genome design, developers may need access to that technology in order to bring acceptable products forward. In that scenario, IP protecting enabling biosecurity technologies could therefore become an important component of the wider commercial landscape.