Biotech has mistaken the accumulation of biological observations for the accumulation of biological knowledge. A company can own more cells, more slides, more patient records, and more perturbation profiles than anyone else and still be unable to answer the only question that creates value: what should we do next?
The failure begins at the measurement boundary. Conventional single-cell RNA sequencing destroys the cell to measure it. The model can learn the response distribution of comparable cells after perturbation, but its output is often interpreted as if it recovered what the measured cell itself would have done. If two biologically different cells look identical to the assay available at decision time yet respond differently to the same drug, more examples under the same assay-and-intervention design can estimate their mixture-average response with exquisite precision. They cannot reveal which hidden condition governs the patient, tissue, or cell in front of us.
The largest cell atlas in the world can still be the wrong experiment.
The durable advantage in biological AI will not come from data volume alone. It will come from recognizing when the available evidence still supports different actions, running the experiment that separates them, and using the result to make a better decision. I will call this experimental epistemic throughput. The phrase matters only if the improvement can later be audited. Until then, it is a claim about the organization rather than a property it has demonstrated.
I argued previously that the phrase “data moat” commits a category error in biology, because the relevant target is decision identifiability: the explanations still compatible with the evidence may disagree about mechanism, but for a declared intervention, outcome, horizon, and cost, they must either recommend the same action or expose the decision as unresolved. Take the recent OrbitAll paper for example, it shows why the learning problem itself has to be constructed around that target.
The OrbitAll preprint does not ask a neural network to infer the full high-level electronic-structure target from molecular geometry alone. It first runs a cheaper semiempirical calculation for each supplied geometry, total charge, total spin, and modeled environment. Spin-resolved Fock and density matrices then encode the converged mean-field electronic structure produced by that solver. In the reported OMol25 model, an equivariant network learns the difference between g-xTB energies and forces and the more expensive density-functional-theory targets. Equivariance hard-codes how the representation and predictions must transform under rotations and translations; delta learning changes the supervised target from the full high-level quantity to the discrepancy left by the cheaper calculation. “Physics” here is not the opposite of data. It is a cheap approximate solver, symmetry, and prior fitted knowledge concentrated into the representation and learning target.
The headline comparisons are not controlled experiments. The 35-fold data figure is a dataset-row ratio, while the 7.5-million-versus-290-million parameter comparison obscures that the compared UMA-S-1.2 checkpoint activates only about six million parameters for a structure and that OrbitAll also pays for the semiempirical calculation. The reported roughly hundredfold cost reduction in the Claisen-rearrangement case compares OrbitAll’s implicit-solvent minimum-energy-path calculation with transition-state-informed umbrella sampling using UMA in explicit solvent—a minimum-energy path against a finite-temperature free-energy profile. They do not prove that physics beats data, but makes a significant point that data alone is not sufficient. That physics and constraints are the key, and generating any such data can be prohibitively expensive: “Unlike language and image models, generating training data of larger chemical systems through density functional theory (DFT) is incredibly expensive.” - Prof Anima Anandkumar.
The more durable result is architectural: in the demonstrated system, OrbitAll changes the representation and supervised target before training, then learns the high-level residual.
In the OrbitAll tests, the molecular variables required by the representation such as geometry, total charge, total spin, and the modeled environment are supplied. Biology often lacks the analogous luxury: the assay may never acquire the variable that separates possible responses, and no representation can recover absent information without another source of variation or a testable assumption.
Consider two cancer cells whose RNA profiles are indistinguishable at the resolution of the assay before kinase inhibition. One may die while the other survives because receptor occupancy, phosphorylation, metabolic load, chromatin accessibility, cell-cycle position, microenvironment, or recent stress history differs. Calling the count vector “cell state” does not make those variables disappear. It hides the experimental decision that excluded them.
The model sees the measurement. The drug acts on the system.
This is why Live-seq was conceptually important. Instead of inferring a same-cell trajectory from populations destroyed at different times, it changed the assay so an initial transcriptomic measurement could be connected to that cell’s later phenotype. The advance was not a larger atlas. It recovered a joint observation that ordinary single-cell sequencing destroys.
The target is not a complete description of life. It is a predictive state that preserves every distinction in the observed history needed to predict declared outcomes under the candidate interventions. The business requirement is weaker: any remaining predictive ambiguity must either leave the preferred action unchanged or expose the decision as unresolved. Both targets are relative to an intervention family, outcome, environment, measurement process, timescale, and the history already observed. A transcriptome can be sufficient for cell-type annotation and insufficient for an acute drug decision. A pathway state that works over minutes can fail across development. We do not need every microscopic variable. We need the distinctions that matter for the prediction and the decision.
Scale can work. Geneformer’s 2026 study found power-law improvements in held-out masked-learning loss and gains on selected downstream tasks, while a separate 2026 *Nature Methods* study found that the tested architectures often saturated after a fraction of the available cells, including on perturbation prediction. The contradiction is only apparent: cell count is not a unit of biological information. Systema showed how a metric can reward the shared shift from control while missing the perturbation-specific effect; Caduceus showed how the right symmetry can let a much smaller model beat larger ones on a long-range variant task. More rows help when the measurement, prior, and target preserve the distinction a decision requires. Otherwise they estimate the wrong average more precisely.
This changes what a biological foundation system must do. In a kinase program, it begins with the action; advance the molecule, change the dose, restrict the population, or stop and then works backward to the distinctions required before that action is defensible. It records what the assay observes and censors, whether target engagement occurred, and which variables were never measured. Missingness cannot silently become zero, a guide barcode cannot silently become a successful intervention, and a dose written in a protocol cannot silently become intracellular exposure. If receptor occupancy, chromatin state, and metabolism all explain the RNA data yet imply different actions, the system must mark the decision unresolved. It may abstain, choose a robust action, or proceed under an explicit risk rule, but it cannot present the choice as identified.
The unresolved decision must then become a physical experiment. The system preserves the rival explanations and the different observations each predicts, generates feasible intervention–readout pairs, discards experiments for which the rivals predict the same observable result, and ranks the remainder by how much each could reduce the cost of a wrong decision under the budget and robustness constraints. It preregisters the outcome thresholds before execution. In the kinase example, that might mean randomized dosing with phosphoproteomic target-engagement measurements, lineage-resolved imaging, or a metabolic readout in the resistant context.
A recommendation earns the right to become an action only by beating the declared baseline, under a fixed budget, in donors, contexts, or laboratories that did not generate the hypothesis. No prediction bypasses the trace: unresolved decision, discriminating experiment, proof that the intervention occurred, out-of-sample result, then action or abstention. If target engagement or assay controls fail, the intervention or measurement contract returns for repair; the execution cannot count as evidence for or against the biological hypothesis. If the prediction fails after valid execution, the state representation reopens instead of absorbing the error through silent fine-tuning. That is the architecture.
Closing the loop still does not make it correct. An acquisition policy uses a model to interpret existing evidence and then uses the same model to decide which evidence will exist next. It can become exquisitely efficient at confirming the wrong ontology. Some experimental budget must therefore be spent against the preferred model: preserving rival explanations, exercising neglected intervention directions, running sentinel experiments the acquisition rule did not choose, reproducing effects across laboratories, and treating systematic residuals as evidence that the state construction is incomplete. The purpose of the loop is to make the model efficiently falsifiable.
LUMI-lab offers an early, bounded example. A pretrained molecular model, active learning, robotic synthesis, and biological testing were coupled across ten cycles and more than 1,700 synthesized lipids. The important object was neither the checkpoint nor the resulting dataset. It was the controlled transition from prediction to synthesis to measurement to update, repeated until an unexpected brominated-tail motif survived physical testing.
A biotech operator can still object that some datasets plainly are moats. The objection is partly correct. Exclusive longitudinal cohorts, linked genotype–phenotype records, rare biospecimens, proprietary perturbation outcomes, and regulatory-grade evidence can be difficult to reproduce. Regeneron’s company announcement about its Truveta collaboration describes a candidate source of defensibility: exclusive research-sequencing rights, consent, and linked health outcomes can create a renewable evidence channel competitors cannot cheaply copy. Privileged access may itself be a commercial moat because it blocks substitutes and creates bargaining power. It is not yet evidence of a scientific learning moat or independently validated economic return; that requires renewable access to produce repeated improvement in real decisions.
The opposite case is equally instructive. Recursion reports more than 50 petabytes of multimodal data and laboratory capacity above two million experiments per week. It has also converted maps into collaboration payments and programs, while the same filing reports a $312.8 million net loss. Those facts do not establish that the platform failed or tell us its scientific learning rate. They show that data volume and laboratory throughput do not, by themselves, demonstrate attractive unit economics; public filings do not expose the conversion well enough to infer more.
A durable moat requires all three conversions to hold. The system must turn uncertainty into a discriminating experiment, the experiment into a better decision, and the better decision into an asset or workflow whose value the organization can retain. Break the first link and the company owns a large archive. Break the second and it owns an impressive laboratory. Break the third and it has built an efficient research service unless its contracts, intellectual-property rights, or ownership structure capture a meaningful share of the value created.
Four quantities must remain separate. Data access measures exclusivity, linkage, coverage, and renewal. Experimental throughput measures quality-controlled experiments per unit of time and total cost. Decision quality measures whether those experiments reduce the cost of choosing an inferior action against a frozen baseline in a new context. Captured value measures what the organization retains after the cost of the platform.
Epistemic throughput is the rate at which quality-controlled experiments improve that third ledger. It must be reported within comparable classes of decisions and against total elapsed time and experimental cost. The available actions, outcome, time horizon, cost of error, baseline, held-out context, and budget must all be fixed in advance. Earlier rejection of wrong hypotheses, fewer go/no-go errors, higher prospective hit rates, and sharper patient selection are consequences of improved decisions, not interchangeable units of one score. A commercial moat requires this scientific capability to be joined to privileged access or control and a credible mechanism for appropriation. Collapse the four ledgers into one platform metric and the result is branding.
Companies that cannot distinguish repeated measurement from resolved uncertainty will spend faster, benchmark better, and still discover less. The next biological platform will not win because its checkpoint has seen the most cells. It will win by finding the ambiguity that changes an action, running the experiment that separates the live possibilities, and retaining the right to compound what it learns.
Organizations that cannot build this loop will not merely own weaker models. They will become more certain about decisions their evidence never identified.
But biology sets a harder limit than strategy. If two biological worlds fit every observation we can feasibly acquire yet demand different actions, what experiment could make the decision identifiable and what should a foundation model do when no such experiment exists?




