In AI drug discovery, the model was never the moat
AI can design the molecules now. The advantage belongs to whoever owns the data.
For most of the last decade, the interesting question about AI in drug discovery was whether it worked at all. Could a model actually design a molecule that a chemist would recognize as sensible, that binds to the target you want, and that a body can tolerate? For a long time nobody really knew the answer.
In 2025 an AI-designed molecule read out positive Phase IIa data for the first time. Insilico Medicine’s rentosertib, a treatment for idiopathic pulmonary fibrosis, is by the company’s account the first drug where the software picked both the target and the compound. The question is no longer whether AI can find drugs, but which companies building this technology will still be around in five years, and which will have been quietly commoditized.
Here is how we think about it.
The economics that make this worth doing
Bringing a single drug to market takes ten to fifteen years and, by standard estimates, north of two billion dollars, and roughly nine out of ten candidates that enter clinical trials fail. The number of new drugs approved per billion dollars of research spending has roughly halved every nine years since 1950, an effect named Eroom’s law, Moore’s law spelled backwards. A process that reliably gets slower and more expensive naturally attracts people who think they can turn it around, and the money has followed.
The AI drug discovery market was worth around $3.6 billion in 2024, and forecasts put it near $50 billion by 2034, or whatever number the latest Wall Street report needs it to be. The partnership money is easier to count. In the last eighteen months Lilly signed with Insilico for up to $2.75 billion, Takeda with Iambic for up to $1.7 billion, Sanofi with Earendil twice for up to $1.72 billion and then up to $2.56 billion.
Treat every one of those numbers with suspicion. Across 516 licensing transactions announced in 2025, upfront cash averaged about seven percent of headline value, roughly fifteen to one, with the rest strung across milestones that, given base rates in this industry, mostly will not be hit. The Lilly deal is the cleanest illustration: $2.75 billion in headlines, $115 million in cash.
Everyone is building a slightly different machine
We have looked at dozens of drug discovery companies over the past few years. They cluster into a handful of distinct approaches, each with its own economics and its own failure mode.
Some design molecules from scratch with generative models, exploring regions of chemical space no human chemist would ever hand-draw. Some skip the target hypothesis entirely and photograph millions of cells to learn what a compound does to living tissue, then let the images tell them what is interesting. Some build enormous knowledge graphs that stitch together papers, trial results, and patents to surface a connection nobody noticed. Others lean on physics, simulating how a molecule binds so precisely that they can save a medicinal chemist months of trial and error. A growing set is wiring AI decision-making directly into robotic labs, so the system proposes an experiment, runs it overnight, reads the result, and designs the next one without a human in the loop.
The archetype does predict durability, but not for the reason most landscape maps imply. It matters because it proxies for one thing: whether the company generates data that did not exist before it ran the experiment.
At one end sit the self-driving labs and the phenomics platforms. Neither can operate without producing proprietary physical measurements as a byproduct, which makes their moat almost impossible to erode with software alone. At the other end sit the pure foundation-model plays, whose principal asset is a set of weights that a better-funded team can match in a couple of months. Generative chemistry and knowledge graphs fall in between, and where any individual company lands depends almost entirely on whether it runs its own experiments or reads someone else’s.
Cutting across all of that is a more commercial question, and it matters more than the technology choice. How does the company make money?
You can sell tools, licensing your software and models to pharma the way a supplier sells picks and shovels to miners. You can develop your own drugs and capture the value in the pipeline. Or you can do both, funding the pipeline with tool revenue while the pipeline validates the tools.
Siemens agreed in 2025 to buy Dotmatics, a scientific software company that never developed a drug of its own, for $5.1 billion, completing the acquisition in 2026. On more than $300 million of revenue, 95 percent of it recurring subscription, that is roughly seventeen times sales for a business with no pipeline at all. A pure tools company can reach real revenue on a few million dollars of seed capital, iterate quickly, and show a customer count that de-risks the next round. At the same time, the easiest way to persuade pharma partners that your platform is working is to have your own assets to prove it.
The pipeline-first path can work when the underlying biology is genuinely differentiated and you have real conviction in it, but there are numerous cautionary tales. BenevolentAI went public at over a $1 billion valuation, its lead drug failed in Phase IIa, and with no licensing revenue to cushion the fall it ended up cutting a large part of its workforce. When the clinical bet is the whole company and the bet loses, there is nothing underneath it.
What just changed, and why it matters most
In early 2026 the largest AI labs walked directly into biology. Anthropic acquired Coefficient Bio for $400 million to build a life sciences division. OpenAI launched GPT-Rosalind, a reasoning model tuned for biology and chemistry, with early access going to Amgen, Moderna, and Thermo Fisher Scientific.
You can read this two ways. The optimistic read is that the buyer pool for startups just expanded overnight. When Anthropic pays $400 million for around ten computational biologists at a company less than a year old, it resets what a data-rich seed company might be worth, and the list of possible acquirers now includes frontier labs and big tech alongside the usual pharma names. The pessimistic read is that the model layer is being commoditized faster than anyone expected. If your entire advantage was that you fine-tuned a model on biological data, you are now competing against some of the most capable and best-funded engineering teams on earth doing the same thing.
My own view is that the big labs are wading into biology partly to improve the story they tell investors ahead of their IPOs. As open-source models keep catching up on performance, the labs need fresh frontiers to justify their valuations, and drug development is a huge, deeply inefficient market where you can gesture at a few more trillion dollars of value waiting to be captured on paper. That makes it a very useful thing to be seen working on. It also makes it easy to abandon. OpenAI shut down its Sora video app and announced the end of its Atlas browser within months of launching each with real fanfare. Those were consumer bets and the biology work looks more considered, but a programme that exists partly to be seen is still a programme that can stop being useful.
Whichever way you read it, the conclusion is the same. As the general-purpose models get better and cheaper, the value drains out of owning a model and pools around owning the data those models need. Better models make good proprietary data more valuable, not less, because a sharper tool makes a rare material worth more. Algorithms are being commoditized. Datasets generated through a unique experiment, a specific assay, a physical lab, cannot be downloaded or replicated by software alone.
This is where we have been putting our own money. DeepSeq.AI runs experiments that return readouts orders of magnitude richer than any public structure database, which lets its models predict the messy properties that decide whether a drug works. Elucidate Bio reads spatial multi-omic data off a single tissue slide, a blend of hardware, reagents, and biology that no model can regenerate on its own. GT Biosciences is assembling an in vivo drug delivery atlas at cell-level resolution. All three sit on the same side of the line.
So the test we now apply to any new company is a simple one. Imagine a world where the best biological AI models are free and available to everyone. Does this company become more valuable in that world, or less?
The practical implication is not that you need a wet lab. It is that somewhere in your loop there has to be a step that produces a measurement nobody else has. A product that gets more valuable as it is used, because usage generates data, compounds.
Where this leaves us
The field spent ten years proving that AI can design a molecule. That battle is won, but many hard problems remain. Translating a computational hit into a drug that works in a human. Generating biological data nobody else has. Building a business that survives a failed trial.
That is the bet. The winners of this vintage will be the companies that manufacture biological data nobody else can, the ones the frontier labs will want to buy rather than rebuild, while the companies whose only real asset was a clever model get quietly commoditized.
Next time I will make the case for the archetype I would back first, the companies generating proprietary human biology to choose better drug targets, since picking the wrong target is still the single biggest reason drugs fail in the clinic.

Jan Buza, Partner at ZAKA VC
ZAKA is an early-stage fund investing in the teams building the next generation of healthcare and life sciences companies.