The exploration budget
When Anthropic launched Claude Science on June 30, the most revealing line in the announcement was what the product wasn’t. Not a new model. Not a more capable model for biology. It runs, in the company’s own words, the same Claude models already available. The pitch was that it removes the friction around the model, connecting to more than sixty scientific databases, coordinating a set of domain sub-agents across genomics and proteomics and cheminformatics, running the analysis in Python and R on the lab’s own cluster, and drawing the figure at the end.
For a company whose entire history is making the model smarter, shipping a science product and saying the intelligence is not the point is a tell. It says the thing slowing science down was never mainly the intelligence. It was the cost of getting each idea from a question to a computed result. And once you take that seriously, science starts to look less like a search for smarter people and more like a problem every investor would recognize.
Science was always an allocation problem
The romantic picture of discovery is the lone mind and the flash of insight. The working reality of a lab is closer to a portfolio under a budget. Ideas are the cheap part. Every field has more plausible hypotheses than anyone can afford to test, and the scarce inputs are mundane: scientist-hours, compute-hours, instrument time, wet-lab runs, the three years of a graduate student’s attention. Every hypothesis a lab pursues is a dozen it does not.
So the real work, the part that never makes the paper, is triage. A principal investigator spends their judgment deciding which few of many candidate directions are worth a scarce slot, and most ideas die there. Not because they were wrong. Because they were not worth the cost of finding out whether they were wrong.
Drug discovery makes the arithmetic stark. The Tufts Center for the Study of Drug Development puts the cost of bringing a new drug to market at around $2.6 billion, over ten to fifteen years, and roughly 90% of the candidates that reach clinical trials fail, most of them in Phase II for simple lack of efficacy. That is not a story about incompetence. It is the base rate of a search problem where you can only afford to buy a handful of the tickets you can see, and you cannot tell in advance which one pays. The budget forces you to concentrate, and concentration in a search this uncertain is mostly a way of buying fewer tickets.
Two ways to change the math
AI does not repeal the budget. What it changes is the cost of a bet, and it changes it in two directions at once.
The first is the one Claude Science is built around: it makes each hypothesis cheaper to chase. The parts of pursuing an idea that used to eat weeks per hypothesis, the literature sweep, the pipeline code, the job on the cluster, the plot for the manuscript, compress toward hours. A generalist agent takes a plain-language request, breaks it into subtasks, and hands them to sub-agents preconfigured for a specific corner of the work. When the cost of carrying an idea from question to result falls, you can afford to carry more ideas on the same budget.
The second is quieter and matters just as much: it makes the losers cheaper to kill. You can narrow the field in silico, computationally, before you spend the genuinely expensive resources on it, the wet-lab run or the scarce compute allocation or the trial. Kill the weak candidates early and cheaply, and the expensive resources flow only to the survivors.
These are the two halves of every search problem. Explore more arms, and prune the bad ones faster. Rationing pushed labs toward exploitation, toward the safe bets that were most likely to return something publishable. Drop the cost of a trial and the optimal strategy tilts back toward exploration, which is where the large, unexpected results live.
Why breadth is the point
Here is the part that an investor sees faster than a scientist might. Discovery pays off like a power law. Most experiments return very little. A rare few remake an entire field, and the winners are worth orders of magnitude more than the also-rans. When payoffs are that skewed, and when you genuinely cannot predict in advance which bet lands, the portfolio logic is not subtle. You want more bets, not a smaller number of more confident ones. Diversification is not a hedge here. It is the strategy.
Which means the concentration that scarcity forced on science was never the clever move. It was the constraint wearing the costume of a strategy. Every hypothesis a lab could not previously afford to chase was a lottery ticket left on the table, in a lottery whose jackpot funds everything downstream of it. Widen the exploration budget and you are not buying a marginal efficiency. You are buying more tickets in the only game where more tickets is close to the whole point.
Concentration was never the clever move. It was the constraint wearing the costume of a strategy.
Read that way, Anthropic betting on workflow rather than a new model is the same insight from the supply side. The marginal smart-scientist-hour was not the scarce input at the frontier. Even at the level of raw capability, one physicist described the current models as about as able as a second-year graduate student at executing a project, which is useful but hardly a shortage that a slightly smarter model would resolve. The scarce input was the number of directions a lab could afford to keep alive at once. Attack that, and you have moved the thing that was actually binding.
The catch is the grading
There is a failure mode here, and it is worth stating plainly, because the optimistic version of this essay is wrong in a specific way.
Widening exploration without improving verification does not give you more discoveries. It gives you more plausible-looking leads, produced faster than anyone can check them. This is the hard problem I have written about elsewhere: scientific code tends to fail silently, returning a clean number and a convincing plot that are quietly, confidently wrong, a sign convention flipped or a method that is unstable in exactly this regime. A wider net catches more fish, and more debris, and at a distance the two look alike.
So the two levers have to move together. Cheaper exploration is only worth something if it is paired with cheaper, more trustworthy pruning. It is not an accident that Claude Science ships a separate fact-checking agent to validate citations and calculations, and leans as hard as it does on reproducibility, results you can trace, figures that carry their own code. That machinery is not polish around the interesting part. It is the load-bearing half. The exploration is worthless if you cannot grade it, and the faster you explore, the more grading you owe.
For most of its history, science rationed curiosity because curiosity was expensive. You could imagine far more than you could ever afford to check, so most of what you imagined stayed imagined. That constraint is the one now coming loose. What replaces it is not a smarter kind of scientist. It is a wider budget for being wrong cheaply, which, in a search this uncertain, is the same thing as a better chance of being right.