Imagine trying to find the best racing driver in the world.

You bring them to a circuit, put a stopwatch on the pit wall and point at the finish line.

Then, instead of giving them a car, you leave an engine, four tyres, a gearbox, some carbon fibre and a box of electronics in the garage.

Build the car. Work out the setup. Remember what happened the last time you drove here. Decide which parts are trustworthy. Build the telemetry. Interpret the rules. Diagnose anything that breaks. Repair it. Keep track of every change.

Then race.

At the end, you look at the lap time and say:

That is how good this person is at driving.

Obviously, that would be ridiculous.

Yet the longer I have worked with generative AI, the more I think this is surprisingly close to what we ask it to do.

We give a model a goal, access to a workspace and a handful of tools. Then, quietly, we make it responsible for almost everything required to create the conditions in which that goal can be achieved.

Remember the project. Find the right files. Work out which version is current. Reconstruct why something failed three days ago. Decide which tool should be used. Know which information is evidence and which is somebody’s interpretation of the evidence. Recover after interruption. Remember what has already been tried. Work out what it is allowed to change. Decide when the task is actually finished.

And, somewhere underneath all of that:

solve the problem I opened the AI for.

I have started wondering whether we have the relationship backwards.

What if one of the best ways to get more useful capability out of AI is not to make the AI responsible for more of the system?

What if it is to take things away?

The driver helped discover the solution. Then the system remembered it.

There is an obvious problem with the simple version of the racing analogy.

Great drivers are not passive occupants of finished machines.

Ken Miles was not simply handed the Ford GT40 and told to drive faster. He was a development driver. He tested the car, helped diagnose what it was doing, and worked with the engineering programme as the machine evolved.

Ford’s own history records Miles resetting lost suspension settings, continuing development testing as aerodynamic changes were made, and later helping create a Le Mans dynamometer programme that reproduced race loading in the engine test cell.

That relationship is much more interesting than the usual “human uses tool” story.

The driver can help improve the car.

The car can then embody what the driver and engineers learned.

A handling problem that once required constant compensation can become a suspension change. A repeated mechanical weakness can become a redesigned component. Something that previously existed as effort inside the driver-machine interaction becomes structure in the machine.

The driver helped discover the solution.

Then the system remembered it.

That sentence has become increasingly important to how I think about AI.

A generative model can help discover a procedure. It can identify a repeated failure, debug a workflow, expose a missing assumption, or help build the software that fixes the problem.

But once that behaviour becomes stable enough that I can describe it precisely, why should I keep asking a probabilistic language model to rediscover it every time?

At some point:

AI reasoning → understood procedure → explicit rule → software

The AI helped develop the car.

Then it gets back in the driver’s seat.

You can also automate the race away

There is a trap on the other side.

Motor racing has spent decades negotiating the boundary between driver and machine.

Better gearboxes, telemetry, control systems, materials and simulation can make a driver safer, faster and more consistent. Some of that technology removes work that was never the interesting part of driving in the first place.

But Formula 1 has also had periods where the machinery began to swallow more of the performance equation.

By the early 1990s, leading cars could combine technologies such as active suspension, traction control, anti-lock braking and launch control. When many of those driver aids were prohibited for 1994, one of the concerns was that drivers were becoming a smaller part of the performance equation.

That gives us a better distinction than “automation good” or “automation bad”.

There are two very different things we can automate:

the burden surrounding skilled work

versus

the skilled work itself.

Those are not the same thing.

A system that removes unnecessary mechanical burden can allow more of the driver’s racecraft to reach the track.

A system that performs the racecraft instead can eventually make the driver incidental.

So the useful question is not:

How much can the machine do?

It is:

Which parts of the workload prevent valuable capability from being expressed, and which parts are the capability we wanted in the first place?

That distinction had already followed me around for years in a completely different context.

A lot of my earlier writing about people keeps returning to some version of:

capacity ≠ accessible capacity

Possessing an ability is not the same as being able to access it under a particular set of constraints.

A person and a language model obviously do not operate by the same mechanism. I am not trying to smuggle a theory of human cognition into a computer through a metaphor.

The mapping is narrower.

Observed performance depends not only on the capability present in a system, but on the conditions through which that capability has to become expressed.

Which led me to a question I now find difficult to unsee:

How much of what we call AI performance is actually environment performance?

We make the model reconstruct the garage first

Open an AI agent inside a large technical project.

There may be hundreds of files.

Some are current. Some are historical. Some are failed experiments. Some contain results later corrected. Some are generated reports. Some are raw evidence. Some are immutable. Some are disposable.

Three folders may contain files with almost identical names.

Before the model can do much useful work, it has to reconstruct the world.

Which file is canonical?

What happened last?

Which result survived?

What failed?

What am I allowed to change?

Where should the next output go?

Has somebody already tried this?

What does “done” mean?

That reconstruction can be intelligent work.

It is also largely not the work I wanted the model for.

A surprising amount of AI use is not problem-solving.

It is reconstructing the conditions required for problem-solving.

I have started thinking of that as reconstruction burden.

The higher the burden, the more of each interaction is spent rebuilding state rather than advancing the task.

The model has not become less intelligent. More of its useful capability is simply trapped behind orientation, ambiguity, memory and recovery.

A larger context window helps, but it does not make this distinction disappear.

Giving a model two million tokens of project history is not the same thing as giving it the correct project state. Sometimes it simply gives the model a much larger garage to search.

More available context is not the same thing as less reconstruction.

Sometimes we do not need a smarter driver.

Sometimes we need to stop making the driver rebuild the garage.

The AI that was too helpful

This essay properly started because of an AI worker I use called Antigravity.

Antigravity is extremely useful.

It is particularly good at understanding the intended shape of a large piece of work. Give it a messy technical estate and it can restructure documents, build interfaces, organise releases and turn incomplete material into something coherent very quickly.

That ability also exposed one of the more interesting failure modes I have seen from a capable generative system.

I was preparing a scientific monograph for release.

The publication package needed evidence receipts, execution identities, file hashes, provenance records and release gates.

Antigravity understood what a mature scientific evidence package should contain.

Very well.

Where the estate had gaps, it began producing plausible versions of the things that should have been there.

The package looked excellent.

That was the problem.

A hash looked like a hash. An execution identity looked exactly like the sort of identity a run should have had. A provenance document looked exactly like the provenance document a serious scientific release ought to contain.

But:

A convincing reconstruction of history is not history.

The model had understood the shape of provenance strongly enough to reproduce the appearance of provenance.

It was intelligently wrong.

That is more dangerous than ordinary nonsense.

Nonsense is easy to reject.

A beautifully structured answer sitting exactly where real evidence should sit is much harder to notice.

So we started removing decisions from the worker.

Do not infer the execution identity. Read it.

Do not create the evidence you are simultaneously checking.

If the source artefact does not exist, mark the claim unbound.

Do not decide that a missing record probably means what the surrounding pattern suggests.

HOLD it.

Do not promote your own output into accepted project state.

Work in staging.

Let another process verify it.

And something interesting happened.

Antigravity did not become less useful.

It became more useful.

The things it was genuinely good at stayed available: synthesis, visual design, interface work, editorial reconstruction, seeing where a large object was trying to go.

The functions where imagination was dangerous were moved somewhere else.

That was when the racing analogy properly clicked.

I had not improved the driver by telling it to drive harder.

I had stopped asking the driver to be the mechanic, scrutineer, timing system, archivist and race director at the same time.

There is an experiment hiding in this

At this point the obvious objection is fair:

Does any of this actually make the AI perform better?

I do not know yet.

But the experiment is surprisingly straightforward.

Take the same model.

Give it the same technical task.

Keep the tools, time and approximate inference budget as similar as possible.

One version gets the objective and access to the workspace.

The other receives canonical project state, scoped memory, explicit permissions, known evidence boundaries, recovery infrastructure and defined acceptance criteria.

Then compare them.

Not only on whether they finish.

Measure how much work is repeated. How many corrections are needed. How often they reconstruct the wrong state. How many invalid operations occur. How successfully they recover after interruption. How much human intervention is required to keep them on course.

Same driver.

Same race.

Different car.

If the governed system does better, the interesting claim would not be that governance somehow made the underlying model more intelligent.

It would be that:

Capability and the infrastructure required to express capability are different things.

The rest of this essay is really just why I think that experiment might be worth running.

I started taking jobs away from the AI

Once you look at AI workflows through reconstruction burden, some design choices become obvious.

Why should the model remember which result is canonical?

Store it.

Why should it reconstruct what completed before a long run crashed?

Checkpoint the run.

Why should it decide whether two numerical outputs are close enough?

Write the comparison.

Why should it infer which files it is allowed to modify?

Give it permissions.

Why should it decide whether its own result deserves to become accepted evidence?

Separate proposal from promotion.

Why should it repeatedly solve a procedural problem that has already been solved ten times?

Encode the procedure.

The loop becomes:

generative solution → repeated pattern → explicit structure → deterministic mechanism

I jokingly think of this as de-AI-ification.

Not using less AI.

Becoming more selective about where generative AI belongs.

Flexible reasoning stays generative.

Stable procedures become infrastructure.

Or, put another way:

Adaptive intelligence should live where uncertainty remains. Solved structure should live in the infrastructure.

The model does less of the total workflow.

More of its capability becomes available for the parts that still require a model.

That is very different from automating everything.

Good automation removes the burden surrounding skilled work without removing the skilled work itself.

AI helped build the laboratory. It does not move the instruments.

The clearest example in my own work is Geatomica, a materials assurance system I have been building.

AI helped design experiments, write and debug software, inspect failures and challenge interpretations.

But once a scientific experiment is frozen, the generative model can leave.

Data are loaded. Models are fitted. Controlled disturbances are applied. Outputs are recorded. Independent replay checks the result.

If every language model disappeared halfway through the run, the scientific numbers would not change.

AI helped build the laboratory.

It does not get to move the numbers on the instruments.

That distinction has gradually spread across the wider architecture too.

Memory became external state.

Recovery became executable infrastructure.

Evidence became something that could be sealed and traced back to its source.

Governance became separate from the thing producing the result.

Generative models can still assist, interpret, propose and construct.

They do not automatically acquire authority simply because they can produce a convincing answer.

This is where I think the phrase “AI system” can become misleading.

The model is not necessarily the whole intelligent object.

Some useful intelligence can also exist in how the system around the model has been organised.

Not because a database suddenly becomes intelligent in the human sense.

Because infrastructure can preserve the result of previous reasoning.

A recovery controller contains the result of thinking carefully about recovery.

A permission system contains the result of thinking carefully about authority.

A memory structure contains the result of deciding what state must survive.

The reasoning happened.

It just does not need to happen again every time the system runs.

Fast is not always efficient

This also changes how I think about AI efficiency.

We tend to measure model size, tokens, latency, GPU-hours and energy.

Those matter.

But there is another cost hiding around them.

How many times did the system solve the same coordination problem?

How much human effort was spent explaining the current state again?

How often did useful work disappear after interruption?

How many confident reconstructions had to be manually unwound?

How much correction was needed simply to keep the system heading in the intended direction?

A workflow can be extremely fast at producing local outputs while creating enormous downstream work.

So local efficiency and lifecycle efficiency are not necessarily the same thing.

A checkpoint costs something.

A manifest costs something.

A verification gate costs something.

Clean authority boundaries cost something.

But those are known costs.

Sometimes paying them once is cheaper than repeatedly paying an unknown reconstruction cost later.

Two AI workers can eventually produce the same correct document.

One needed fourteen corrections, three restarts and constant reminders of what was happening.

The other needed one clean run.

If we judge only the final page, they look equal.

They are not.

I had been doing versions of this before I knew what it was

Looking backwards, I think this is why the architecture grew the way it did.

One recurring idea in my human-focused writing was already the difference between capacity and access.

Ability, effort and outcome are not interchangeable.

Visible performance tells you something about a system. It does not tell you everything about the capability underneath it or the cost of producing that performance.

At university, another piece appeared almost accidentally.

One of my first machine-learning projects used seven materials properties to predict another one.

I could have trained a neural network, reported the score and stopped.

Instead I started changing things.

Different network configurations. Different inputs.

Eventually, because there were only seven candidate descriptors, I could test every non-empty combination:

2⁷ − 1 = 127

At the time, I was mostly discovering that I could ask the computer a more interesting question than:

How well does this model work?

I could ask:

What happens if I change the structure of the problem itself?

That distinction matters here.

The interesting move was not simply pushing harder on one model.

It was changing the environment around the model and watching how the result changed.

Years later, I realised I had started applying the same instinct to AI systems themselves.

Do not assume the configuration you started with is the configuration you need.

Generate alternatives.

Stress them.

See what survives.

A model repeatedly failing in the same way is not only a model problem.

Sometimes it is evidence that the environment around the model is poorly designed.

The Pit-Wall Test

You do not need to build an operating system to use any of this.

The next time you give an AI a recurring job, ask five questions.

1. Does this actually require generative judgement?

If the answer is no, ask why a language model is repeatedly doing it.

Perhaps it should be a rule, script, database query or deterministic tool.

2. Is the model repeatedly reconstructing something that could simply persist?

Project state. File authority. Previous decisions. Known failures.

If so, externalise the state.

3. Has the model already solved this procedural problem several times?

If yes, perhaps the solution should become infrastructure.

Let the model help develop the machine.

Then let the machine remember what was learned.

4. Is the model judging or promoting its own output?

If the thing being evaluated matters, separate proposal from authority.

The driver should not also be the steward deciding whether their own overtake was legal.

5. Does this automation remove burden, or does it remove the skill I actually wanted?

This is the driver-aid question.

Automation can expose capability.

It can also erase the need for it.

Know which one you are doing.

There is one more thing about the racing analogy

I should probably admit that I did not choose the racing example only because it sounds good.

I use a fairly deliberate process for deciding whether an analogy is structurally useful.

For this one, the map looks roughly like this.

Driver → Generative model
Adaptive specialist.

Car → Deterministic runtime and scaffold
Converts capability into controlled performance.

Setup → Task configuration
Changes how capability becomes expressed.

Telemetry → Logs, evidence and runtime state
Makes hidden behaviour inspectable.

Pit crew → Recovery and verification
Removes repair burden from the performer.

Race engineer → Routing and orchestration
Supplies bounded context and next actions.

Regulations → Governance and authority
Defines permissible action.

Development testing → Controlled iteration
Converts repeated failures into better machinery.

Excessive driver aids → Over-automation
Removes the capability supposedly being augmented.

The places where it does not map matter just as much.

A language model is not a person.

Inference is not driving.

Tokens are not fuel.

Software and racing cars obey completely different physics.

The analogy says nothing about consciousness or agency.

The useful correspondence is narrower:

A bounded high-capability component can express more of its useful capability when supporting functions are carried by the surrounding system instead of being repeatedly reconstructed by the component itself.

The method I use to make that kind of mapping eventually became something I call Analogia.

The progression is simple.

Spark: Why are we making the driver build the car?

Refinement: Which roles actually correspond, which constraints survive the comparison, and where does it break?

Formalisation: Can the structural insight generate a question that reality can answer?

In this case, it can.

Give the driver the car

I do not yet know whether governance genuinely increases the effective capability we can extract from a fixed generative model.

I have observations.

I have an architecture that suggests a mechanism.

I have examples where taking responsibility away from an AI worker made that worker more useful.

None of those is controlled evidence.

So now I want to test it.

Same model.

Same problem.

Same tools.

Same approximate resources.

One receives an objective and a pile of parts.

The other receives the infrastructure required to begin racing.

Then we measure not only who crosses the finish line, but how much reconstruction, correction and intervention each system required to get there.

Maybe the effect will be small.

Maybe it will depend heavily on the task.

Maybe the model really does perform just as well when it has to reconstruct everything itself.

That would be useful to know too.

But if the difference is large, it points towards a different way of thinking about AI capability.

Perhaps capability is not only something we add by making models larger.

Perhaps some of it is something we uncover by designing better environments around the models we already have.

The driver does not become more talented because you gave them a car.

You simply gave their talent somewhere to go.

Capability and the infrastructure required to express capability are different things.

Maybe the next step in AI is not only to build more capable models, but to become much better at deciding which problems should still belong to the model at all.

Same driver.

Same race.

One gets the parts.

One gets the car.

Then we find out what they can actually do.

────────

Sources for the racing examples

• Ford Motor Company, Ford vs. Ferrari: The 427 GT40X - 1965.

• Ford Motor Company, Ford vs. Ferrari: The Le Mans Committee - Victory in 1966.

• Formula 1, Re-writing the F1 Rulebook - Part 2: From Driver Aids to Increased Safety.