I. The bottleneck in modelling has quietly changed.
For much of the history of computational modelling, predictive accuracy has been the defining measure of progress. Better models were difficult to construct, expensive to evaluate, and often represented years of incremental advances in mathematics, computing and experimental understanding. Under those conditions it made perfect sense to judge progress primarily by predictive performance. If one model predicted reality more accurately than another, it was usually the more valuable scientific instrument. Improving prediction meant improving science. That assumption is beginning to change. Across engineering, materials science, biology, finance and increasingly every discipline touched by computation, we are entering an era in which producing another highly accurate predictive model is no longer the scarce resource it once was. Machine learning, automated model discovery, foundation models and rapidly expanding computational capability have fundamentally altered the economics of prediction. Tasks that once demanded months of specialist effort can now be explored in hours. Entire families of candidate models can be generated, optimised and compared automatically. In many domains, prediction is becoming increasingly abundant. This is not because modelling has stopped improving. Quite the opposite. Modern models routinely discover statistical relationships that would have been inaccessible only a decade ago. They interpolate enormous datasets, uncover subtle interactions and increasingly contribute to scientific discovery itself. Their success, however, has quietly exposed a different limitation. As predictive performance becomes easier to obtain, its value as the sole measure of scientific progress inevitably begins to diminish. Accuracy is becoming abundant. What remains scarce is something else entirely. Not prediction. Not optimisation.
Not computation. The ability to determine which predictions deserve our trust. The title of this essay deliberately uses the word commodity, because commodities are not things that have lost their value. Steel is still valuable. Electricity is still valuable. Concrete remains indispensable to modern civilisation. What changes is not their usefulness but their scarcity. As production becomes cheaper and more routine, value migrates elsewhere in the system. I believe predictive modelling is beginning to undergo the same transition. The cost of generating another accurate model is falling rapidly. The cost of determining whether that model deserves engineering trust is not. Every mature technology experiences a shift like this. Steam engines eventually stopped being constrained by steam itself and became constrained by metallurgy. Software development ceased being limited by our ability to write code and became limited by testing, version control and reproducibility. Manufacturing stopped asking whether parts could be produced and began asking whether they could be produced consistently. As production becomes easier, assurance becomes harder. The bottleneck moves from creating outputs to governing them. Computational modelling appears to be reaching exactly the same point. For decades, the central question was simple: Can we build models that predict reality more accurately? Increasingly, the answer is yes. Modern workflows can produce highly performant models with remarkable speed, often across problems that previously required years of manual effort. Yet that very success has created a new question, one that is arguably more difficult than the first. Which of those increasingly accurate models actually deserve our trust? Prediction and trust are not the same thing. Prediction estimates what may happen under a particular set of assumptions. Trust determines what we are willing to believe, investigate, manufacture, publish or spend millions validating. Those decisions carry consequences far beyond a benchmark score. Research programmes are prioritised. Experiments are funded. Components are manufactured. Materials are selected. Clinical studies are advanced. Entire engineering programmes change direction because a model appears sufficiently convincing. Yet the statistical performance of a model and the justification for acting upon it have never been the same quantity. We have simply become accustomed to treating them as though they were. I suspect that the next decade of modelling will not be defined by another dramatic leap in predictive capability. Those improvements will undoubtedly continue, but they are no longer the only story. The more profound transition is quieter. Prediction is becoming increasingly automated. Judgement has not. The scarce resource is shifting from generating another answer to determining which answers deserve to become accepted knowledge.
That, I believe, is where the next bottleneck in modelling quietly waits.
Figure 1. The Bottleneck Has Moved For decades the primary challenge in computational modelling was producing accurate predictions. As predictive capability becomes increasingly abundant, the limiting resource shifts towards determining which predictions deserve engineering trust.
I. Accuracy Was Never the Whole Story
If this argument sounds provocative, it shouldn’t. The claim is not that predictive accuracy has stopped mattering. It remains one of the most important properties any scientific model can possess. Without predictive capability there is nothing to validate, nothing to interpret and nothing to apply. Accuracy is the foundation upon which every other judgement rests. The problem is that it has gradually become treated as though it were the entire building. Two models may achieve almost identical predictive performance while possessing completely different scientific value. One may continue to reach the same conclusions after modest perturbations to the data, modelling assumptions or feature selection. The other may collapse as soon as those conditions change. Both may report the same R². Both may satisfy the same benchmark. Yet only one represents a relationship that appears robust enough to support further engineering effort. Imagine, for example, screening ten thousand candidate alloys for a lightweight aerospace component. Two independent modelling workflows identify the same material as a promising candidate. Both achieve similarly impressive predictive performance against historical data. At first glance they appear equally persuasive. Yet after modest perturbations—slight changes to the training data, alternative feature selections, different modelling assumptions or independent validation—one continues pointing towards the same alloy while the other begins recommending entirely different candidates. Both models remain accurate according to conventional benchmarks. Only one remains reliable enough to justify the next stage of expensive physical validation. The benchmark score did not reveal that distinction. The engineering workflow did. Traditional performance metrics struggle to distinguish between situations like these because they were never designed to answer that question. Metrics such as R², RMSE, MAE or classification accuracy describe how well a model performed under a particular evaluation procedure. They tell us remarkably little about how confidently that result should influence engineering decisions beyond it. This distinction has existed for decades. It has simply become more visible as prediction has become easier. When accurate models were rare, the practical objective was to obtain one. As accurate models become plentiful, comparison becomes inevitable. Once dozens of models can produce similarly impressive scores, the question naturally changes from Which model predicts best? to Which prediction deserves to shape our understanding of reality?
That transition is exactly what we would expect if predictive performance is gradually becoming a commodity. The benchmark remains essential. It simply ceases to be sufficient as the primary differentiator. Scientific maturity increasingly depends not on our ability to generate another prediction, but on our ability to determine what that prediction actually means. That is not merely a statistical question. It is an engineering question.






