Materials ML case study / current public state

Better than null is not yet useful

A model can learn more than a deliberately broken route and still fail a minimal held-out predictive boundary.

Current stateDEMONSTRATED IN A BOUNDED ROUTE
The practical question

Should this model be allowed to promote a candidate into more expensive engineering work?

First ask whether the valid model beats a paired target-permutation null. Then ask, on the same cases, whether it improves on the constant reference represented by positive held-out R2.

Beat paired broken-target null211 / 288
Positive held-out R213 / 288
Passed both13 / 288
Same-denominator comparison: 211 of 288 beat the paired null, while 13 of 288 had positive held-out R squared
Measured on the same 288 Materials Project elasticity model-evaluation cells.

What the observation protects against

Relative success and absolute viability are retained as separate gates.

Observation

Non-null signal was common. Positive held-out R2 was rare.

Limit

R2 greater than zero is only a minimal metric boundary, not a universal engineering acceptance threshold.

Next investigation

Inspect target, population, representation and failure distribution before selecting further simulation, synthesis or testing.

This is public scientific evidence from one bounded materials workflow. It does not establish customer savings, independent validation or deployment readiness.