Act I — Just About Doing It
I have been getting away with things for quite a long time.
Not crimes, generally.
More like functioning.
There is a particular kind of functioning where, from the outside, the relevant thing happened, so nobody has much reason to investigate how close it came to not happening.
You passed the exam.
You handed the thing in.
You turned up.
You answered the question.
Lovely.
Job done.
Nobody usually follows that with:
“Cool. And was the machinery producing this result in anything approaching a sensible condition?”
Why would they?
Most of the time the result is the bit people need.
I think I learnt quite early that if you can keep getting the result through, people will quite reasonably assume the system underneath it is basically alright.
I also learnt that I was unusually good at getting the result through.
This would become both useful and deeply inconvenient.
I grew up around Somers Town in Camden, moving between different houses and different parts of my family rather than having one perfectly consistent little base camp. Home itself was loving. My parents were young, split early, and both stayed in my life. Mum was dealing with disability while also doing a fairly heroic amount of dragging our circumstances upwards. Dad remained present. There were people around me.
Outside that, the environment changed quite a lot.
Schools.
Houses.
Rooms.
People.
Different versions of what was expected depending on where I was.
I don’t remember responding to any of this with a conscious little child manifesto about adaptability.
You just clock things.
Who likes who.
Who is annoyed.
Who can be joked with.
Who cannot.
Whether this room wants loud Joe or quiet Joe.
Whether being clever is currently useful or whether perhaps we keep that one in the chamber.
Children are doing social modelling long before anybody gives them the words for it.
Mine just got quite a lot of training data.
Primary school was socially rough in ways that felt completely ordinary at the time. There was bullying, physical stuff, friendships that could somehow contain both genuine closeness and somebody launching an object at you, and the sort of miniature hierarchy children create when left to invent civilisation from scratch.
I was also doing very well academically.
This did not grant the diplomatic immunity I might reasonably have expected.
I could answer questions aimed at older children and was put into the various “Able, Gifted and Talented” categories schools used back then.
I remain slightly suspicious of the phrase gifted.
It sounds like somebody handed you something useful.
What they actually gave me was a certificate and higher expectations.
Still, the academic bit was real.
A Year 3 report I found recently is almost offensively on the nose. It says I was good at making connections across my learning, enjoyed scientific investigation, could identify what data needed collecting, and understood why a test was fair or unfair. Then, having spent several paragraphs describing a child who apparently wanted to interrogate everything in sight, my teacher gave me two targets.
Slow down and think about the reader.
And explain the mental method used to reach the answer.
Basically:
show your working, mate.
I was eight.
Adult me has taken this feedback to an unreasonable extreme.
But the important bit is not that my teacher accidentally predicted a research programme.
She obviously didn’t.
It is that the mismatch was already visible in a harmless form.
Sometimes I got to the structure before I had built the route another person could follow.
I knew the answer.
Now somebody wanted the corridor plan.
Fair request.
This became a recurring offence.
As school got more socially complicated, another separation became more useful.
What was happening inside me and what was visible outside did not always need to match.
Again, I did not sit down and design this.
I just got good at reading what an environment required and supplying roughly that.
Secondary school turned the difficulty up a fair bit.
I went to an all-boys school where social hierarchy had somehow acquired additional funding.
Reputation mattered.
I got mine wrong early.
Once people have decided who you are, there is a slightly annoying tendency for that version of you to gain persistence.
By the later years, different social groups and schools had started mixing, stories travelled, relationships got messy, and I discovered that information about you can move around perfectly well without requiring your participation.
Then you meet somebody new and, after a few weeks, they say:
“You’re actually nothing like what I heard.”
Which is nice.
At first.
After enough repetitions, you do begin wondering who this other Joseph is and whether he could maybe start contributing financially.
What I got very good at during those years was managing the visible layer.
Read the room.
Adjust.
Do not make the situation worse.
Keep moving.
Deal with whatever is going on internally somewhere less public.
There were periods where that became much more serious than I want this essay to get into, but the pattern matters more than the individual events anyway.
The basic adaptation was:
contain first, process later.
Later is a fantastic concept.
Massive fan of later.
Later has also historically received a lot of work without corresponding increases in staffing.
By the time I reached my mid-teens, the same split had started appearing academically.
There is one memory I keep coming back to because it is so stupidly clean.
I was sitting in front of a maths textbook.
I understood the maths.
That was not the problem.
The problem was that I could not make myself start studying it.
Not “I didn’t fancy it”.
Not “I got distracted after ten minutes”.
I mean sitting there, looking at the thing I wanted to do, knowing how to do it, and somehow being completely unable to cross the tiny invisible gap between intention and action.
Eventually I cried.
Which is quite an aggressive response to a textbook you already understand.
At the time, I had no useful explanation for this.
So I went with the obvious one.
Lazy.
Poor discipline.
Need to try harder.
Maybe just fundamentally a bit of a wasteman.
Standard teenage diagnostic framework.
The difficulty was that this explanation kept running into contradictory evidence.
Because when something became urgent enough, I could suddenly do loads.
A deadline would get close.
The internal shutters would open.
Five hours of work would materialise out of nowhere.
Sometimes it would be genuinely good.
Then the result came back decent and the obvious conclusion was:
See?
Could have done it all along.
Which is technically true in the same way that a car capable of doing 120 mph is therefore perfectly suited to travelling everywhere at 120 mph.
Capacity was there.
Control over access to it was the questionable part.
I did not understand that distinction yet.
So I kept using peak output as evidence of normal capacity.
Other people did too.
That is nobody’s fault, really.
Peak output is very convincing.
If somebody repeatedly pulls things out of the bag at the last minute, eventually you stop asking why everything keeps ending up in the bag.
College and then university mostly extended this arrangement.
COVID did not help, but I was hardly running a textbook educational workflow beforehand.
Online learning was especially dead for me.
Sit in room.
Open laptop.
Watch lecture.
Remain there.
Apparently four separate specialist skills.
When university became more normal again, my attendance did not exactly undergo a triumphant recovery.
My total attendance across the degree is low enough that putting an exact number in writing feels legally unwise.
Let’s say the university and I had a flexible interpretation of presence.
And yet the actual materials science could click very quickly.
That was the confusing bit.
I could miss a ridiculous amount, encounter something relatively late, see how the pieces fitted together and then explain it to somebody else.
At points I relied heavily on people around me for the bits I was bad at.
Structure.
Body-doubling.
Actually turning up.
Knowing what had happened in the lecture I had, through a complex sequence of events, not attended.
Then when we were sitting together with the technical problem in front of us, the balance could flip.
I could suddenly be the one explaining the mechanism.
There is something very strange about being simultaneously good at a subject and bad at being a student of the subject.
Those are not supposed to feel like separate occupations.
For me they often did.
And because I continued getting results, the arrangement remained surprisingly difficult to interrogate.
This is where “just about doing it” becomes a bit of a trap.
Enough output to prove capacity.
Enough dysfunction to make producing it increasingly horrible.
Never quite enough visible failure for the whole setup to become obviously non-viable.
Every time I pulled it off, yesterday’s emergency performance quietly became tomorrow’s baseline.
The day after I submitted my dissertation, I had my ADHD assessment.
The timing remains quite funny to me.
I had essentially completed most of a degree using a combination of panic, pattern recognition, social scaffolding, occasional dissociation and whatever dark magic causes an assignment to become possible at 2am.
Then, the next day, somebody professionally informed me:
“Yeah mate, there may be a reason for that.”
Combined-type ADHD.
One assessment.
Very little suspense.
In hindsight, there had been approximately four thousand clues.
At the time, though, the diagnosis changed something quite important.
It did not suddenly make me more capable.
It did not give me a new personality.
It did not retroactively turn every bad decision into a symptom.
What it gave me was a better question.
For years, when I could understand something but could not make myself do the thing required to act on that understanding, the question had been:
What is wrong with me?
Now it could become:
What is blocking the route?
That sounds like a small change.
It is not.
The first question makes the person the failure.
The second makes you inspect the mechanism.
And mechanisms can be worked with.
Sometimes.
I should be careful here because it would be very easy to tidy the whole story up retrospectively.
Child struggles.
Child masks.
Teenager cannot study.
Young adult gets ADHD diagnosis.
Everything finally makes sense.
Roll credits.
Absolutely not.
I was still improvising constantly.
I still had years of behaviour I did not really understand.
The diagnosis explained one layer of the problem. It did not reveal the full source code.
But it did help me separate a few things I had previously been treating as one.
Knowing something was not the same as being able to act on it.
Being able to act once was not the same as having reliable access to that ability.
Producing a result was not the same as producing it sustainably.
And looking composed while doing all of that told you surprisingly little about how expensive it had been.
I did not yet have names for most of those distinctions.
I definitely did not have a framework.
Mostly I had an increasingly elaborate collection of ways to just about keep things moving.
And to be fair to them, they worked.
That is important.
Masking worked.
Urgency worked.
Being funny worked.
Reading people worked.
Last-minute pressure worked.
Other people providing structure worked.
Compartmentalising worked.
At times, emotionally switching off worked.
I do not think the useful lesson is that all coping mechanisms are secretly toxic and everybody should immediately become maximally authentic in every environment.
That sounds exhausting.
Adaptations exist because they solve something.
Mine solved quite a lot.
The problem was that I had gradually built a life around adaptations designed for environments where there was still enough spare capacity to pay for them afterwards.
For years, there usually was.
Maybe not loads.
But enough.
Then, in 2023, I went to Bristol for my industrial placement and something interesting happened.
Professionally, things got easier.
Not because the work was simple. Almost the opposite.
For once, I was in an environment where wanting to know what sat underneath the visible result was pointed at the right kind of problem.
Here was a material.
Here was a method.
Here was something that had to be made observable properly before anybody could tell the clever story about what it meant.
I liked that immediately.
I did not know it yet, but the placement was about to give me a much more disciplined version of a question I had been circling for years:
what does the visible output fail to tell you about the thing underneath it?
Meanwhile, my own body was quietly becoming a less theoretical version of the same problem.
And this time, “just about” was not going to scale.
Act II — The Component on the Table
Bristol was probably the first place where one of my more annoying habits became professionally useful.
I had spent years wanting to know what was happening underneath things. Why this result? Why that pattern? What had to happen before this thing ended up looking like this? At school this could be a slightly irritating personality trait. Materials engineering is much more accommodating. Sometimes somebody literally hands you a bit of metal and asks you to find out what is inside it.
In 2023 I moved to Bristol for my industrial placement with Intertek. The Intertek team operated as a consultancy office inside the Rolls-Royce Defence Aerospace building, so I was physically working in that environment while carrying out independent materials assessment work for the programme. I arrived not really knowing Bristol, bounced through temporary accommodation for a bit, eventually ended up in a studio with a commute long enough to develop opinions about bus routes I had no previous emotional investment in, and then basically got on with it.
Work clicked surprisingly quickly.
Within the first couple of weeks I was being trusted with serious materials assessment work, including additive-manufactured titanium components. The job was not failure analysis in the full sense. That distinction matters. I was doing most of the physical and measurement work required before somebody could make the final analysis: sectioning samples, mounting them, grinding, polishing, etching, getting them under the microscope, measuring porosity, lack of fusion, grain structures, coatings and whatever else the assessment required, then producing the information somebody else would use to interpret what had happened.
Which, if I am being completely fair, occasionally annoyed me.
Because the bit I naturally wanted was the bit after the measurements.
I wanted to look at the pattern and ask why.
What process produced this? Why is the porosity concentrated there? What does that grain structure tell us? What changed upstream? What explanation fits all of the evidence rather than just the most obvious feature?
Instead, a lot of the time my job was essentially:
Here is the thing.
Make the hidden structure visible.
Measure it properly.
Do not fuck up the evidence.
Pass it on.
There are worse problems to have.
And it taught me something important anyway, because you very quickly realise how much work happens before anybody gets to tell the clever story.
A material does not arrive beneath the microscope conveniently prepared to explain itself. Somebody has to choose where to cut it. The sample has to survive preparation. The surface has to be good enough to reveal the structure you care about. The etch has to actually expose the relevant features. The measurements have to mean the same thing from one sample to the next. If any of that is sloppy, the analysis downstream can be beautifully reasoned around evidence you have already distorted.
That lodged somewhere in my head.
Before interpretation comes making the thing observable properly.
And the thing you observe is still not the process that produced it.
Under a microscope, an ordinary-looking bit of metal becomes terrain. Grains, boundaries, pores, inclusions, regions that fused properly, regions that did not, coatings behaving differently across an interface. The manufacturing process is over by the time you see any of this. Whatever heat, cooling, pressure, geometry or processing variation produced the structure has already happened.
You get what survived.
Then somebody works backwards.
That gap fascinated me.
A pore is definitely a pore. That does not mean the pore has kindly supplied its own cause.
A strange microstructure is real. Your first explanation for it may still be bollocks.
Two materials can look compositionally similar and behave differently because their histories are different. Processing has left something behind in the structure.
I wanted to read that history.
Mostly, my job was to make sure whoever did read it had something trustworthy to read.
There was another part of the placement that stuck with me for a completely different reason. Some technical knowledge was much more fragile than the material in front of us. A coating-assessment process had lost some of the expertise that previously held it together, so parts of the method had to be made explicit again. What exactly are we measuring? How are we calculating it? What does the report need to contain? How do you make the judgement reproducible enough that the next person is not reconstructing the process from vibes and the ghost of whoever used to know it?
I helped build that workflow out, document it and train other people to use it.
At the time I was not thinking about “externalised procedural memory” or any of the other phrases I would later invent after being allowed near too many documents.
It was more basic than that.
If a process only works because one person remembers how it works, you do not really have a process.
You have a person.
Write enough of it down, make the measurements and calculations explicit, leave a route another competent person can follow, and something changes. Knowledge that used to live mainly in somebody’s head starts living partly in the environment around them.
I would reinvent that lesson several times later.
The placement also gave me a much cleaner relationship with ability than university had. Education had always made competence feel strangely conditional. Could I attend consistently? Could I initiate the work? Could I make myself revise something I already understood? Could I perform the required routine for long enough to arrive at the interesting problem?
Work was simpler in that respect.
Here is the component.
Here is the preparation route.
Here is the microscope.
Here is what we need measured.
Go.
My brain likes that.
Give it a difficult physical problem and there is a decent chance it wakes up. Give it a portal requiring three forms, two passwords and an email chain called RE: RE: Important, and suddenly we are negotiating with a separatist region.
So professionally, I was doing well.
Outside work, things were becoming less tidy.
I had gone to Bristol trying to build some structure around myself too. Gym, football, volleyball, cooking, routine. Then I tore my quad, which was rude but manageable. Sport stopped. I adapted. Rehab, different routine, carry on.
That was what I knew how to do.
Something changes, move the load somewhere else.
Route blocked, find another route.
Keep the important bits running.
For most of my life that had worked well enough that I did not have much reason to question it.
Then back pain started. At first it was just back pain, the sort of thing that barely qualifies as a plot development. Then pain spread. Fatigue became less temporary. Recovery started taking longer. The amount of effort required to keep ordinary things ordinary was creeping upwards.
There was not one clean moment where my body announced a regime change.
That would have been convenient.
It was much more gradual. A bit more pain here. Less recovery there. One activity quietly disappears. Something that used to cost ten units now costs fifteen. You compensate, so the visible result remains surprisingly similar.
That is probably why I did not take it particularly seriously at first.
I was still working.
And this is where I managed to spend my days around materials and somehow avoid applying even the most obvious lesson to myself.
Current function is not the same thing as remaining margin.
A component can still carry load while damage accumulates.
That does not mean it is already failed. It also does not mean the current state tells you everything you need to know about what happens next.
At work, this made perfect sense.
At home, my diagnostic framework remained:
Still standing, innit.
Carry on.
There were other pressures outside work during that period too. Some were significant, but I do not think this essay needs to excavate every detail. The important thing is that recovery stopped behaving like recovery. Work remained structured and predictable, so I protected it.
Again, the output was real.
I was actually doing the work.
The mistake was using the existence of the work as evidence that the whole arrangement remained healthy.
There is a seductive little chain there:
I am still working.
Therefore I can still work.
Therefore this way of working is still viable.
Those are three different claims pretending to be one sentence.
I finished the placement.
That matters too.
I went back to university. For a while, medication and lower surrounding load gave me more cognitive access again, and some of my strongest technical work reappeared with it.
The capability was still there.
Very reassuring.
Also a terrible incentive to keep believing the surrounding system could probably take another hit.
Toward the end of 2024, that story became harder to tell. Physical symptoms worsened. Sleep deteriorated. Medication continuity became messy. Autonomic symptoms became harder to dismiss. Eventually there was a collapse serious enough that university stopped and I took medical leave.
By then, the problem had changed.
I was no longer dealing with one thing that needed solving. Pain, fatigue, medication, sleep, heart rate, food, appointments, university, money and mental health were interacting while my ability to organise any of them was getting worse.
The component was no longer sitting neatly on the table.
I was inside it.
And the really annoying part was that I could still see enough of the structure to know there was a bigger pattern.
I just could not reliably hold it still long enough to explain it.
Act III — Keeping the Whole Thing Together
By autumn 2025, my health was already well beyond “a bit rough”.
I had gone back to university and was trying to remain a student, keep up with a group design project, deal with admin, and generally perform the ancient ritual of pretending there were enough hours and functional organs available for all of this. At the same time, I kept ending up back in medical settings often enough that I no longer trust myself to give you a clean count.
That period is hazy.
Not mysteriously hazy. Just the normal kind of haze you get when too much is happening for too long and your brain eventually decides that chronological indexing is a luxury feature.
Pain was already part of everyday life. Sleep was unreliable. Eating had become weirdly optional despite being, biologically speaking, quite strongly recommended. Standing up increasingly came with visual effects. My heart rate was freestyling like prime Blind Fury. If you know, you know.
Then my head started doing something new.
I kept describing it as pressure because that was the only word that felt remotely right.
Not the sharpness I associated with a migraine. Pressure behind and around my eyes, across my face, sometimes with vision changes and dizziness. At its worst, it felt like there was a tiny black hole sitting somewhere behind my skull trying to pull everything towards it.
I knew, rationally, that my face was not being mechanically crushed inwards.
This information was not especially comforting at the time.
The working answer remained migraine. I tried migraine medication. It did basically fuck all.
What did help, slightly, was staying calm.
Not in the wellness-influencer sense. I was not breathing deeply and manifesting a regulated autonomic nervous system. Getting worked up simply seemed to make the whole experience harder to tolerate. So when the pressure became properly diabolical, part of dealing with it was trying not to react to how diabolical it was.
At some point I changed my phone lock screen to KEEP CALM AND CARRY ON.
Which is probably the most aggressively British decision I have ever made.
Given the general diversity of where and how I grew up, I find it quite funny that under sufficient physiological pressure my final emergency protocol was apparently wartime government graphic design.
But there is also a more important explanation.
Begrudgingly, I am British.
I do not want to be a hassle.
This principle survives surprisingly far into medical crisis.
Even the physical-health letter I eventually sent my GP, after several pages describing recurrent near-blackouts, visual problems, neurological symptoms, sleep issues, nerve pain, poor food intake and everything else, ends by reassuring them that I was “not asking for everything to be handled in one appointment.”
Body producing increasingly experimental behaviour.
Still don’t want to cause a fuss.
There were also more advanced clinical interventions available.
Sometimes, when the pain was particularly long, I would put on the deepest voice I could manage and announce:
“Pain? What pain? Pain is but a concept made for the weak to justify their screams.”
Zero clinical evidence.
Strong patient satisfaction.
It did absolutely nothing to the pain. But for a few seconds it changed the relationship to it. The pain became something ridiculous enough to mock rather than the entire environment I was trapped inside.
Then the speech ended.
Still hurt.
Bosh.
Carry on.
The problem with “carry on” was that by then I was carrying quite a lot.
A bad night changed the pain. Pain changed movement. Both changed whether I ate. Not eating changed cognition. Cognitive load changed whether I could organise medication, appointments, coursework or basic life maintenance. Stress could amplify physical symptoms, then the physical symptoms gave me more things to be stressed about.
And ADHD remained impressively committed to its own agenda.
There could be six objectively pressing problems in front of me and my brain would still spot some technical question in the corner and go:
rah, what’s that though?
Three hours later I might understand something genuinely interesting about a neural network and still not have eaten.
Useful brain.
Questionable management structure.
By October, I had stopped believing the problem could be represented as a sensible list.
I had also stopped believing that simply describing each individual symptom more clearly would necessarily fix that.
So I tried to make the relationships visible.
On 16 October I sent my GP a seven-page combined health summary. Right at the top I explained that I had used ChatGPT to consolidate my own words because sustained writing and thinking were worsening the head symptoms. The document then tried to hold physical deterioration, nutrition, sleep, executive function, emotional regulation, healthcare barriers and cognitive overload in one place. There was even a table showing how one problem fed another. Naturally. Give me enough distress and apparently eventually I make a table.
At one point I described myself as trapped in a state of “forced continuation”.
That was probably the closest phrase I had.
I knew carrying on was making things worse.
I also did not seem to possess a reliable mechanism for stopping, stabilising everything, and then sensibly restarting life once the maintenance work was complete.
The plane was being repaired while flying.
The mechanic was also the passenger.
And the pilot had ADHD.
By the time I wrote that letter I estimated I had spent around fifteen hours and roughly 35,000 words talking through what was happening with ChatGPT.
Thirty-five thousand words.
To explain that cognitive load was becoming a problem.
Efficient.
But something important had changed.
I was no longer mainly using AI because it could answer questions.
I was using it because the previous state of the problem could still be there when I came back.
This happened.
Actually, that started earlier.
No, wait, the medication changed here.
This gets worse when that happens.
I thought those two were connected but now I’m less sure.
Put that over there for a second.
Remind me what we already established.
Now help me do the university thing.
Now the group project.
Now this email.
Now back to the health thing because my body has released a patch nobody requested.
I already had early Master Scripts and structured ways of getting the model to hold complicated thought. But during this period they stopped being merely interesting intellectual machinery.
They were being conscripted into running my life.
Health.
Degree.
Design work.
Admin.
Ideas.
Chronology.
Things I had to tell a doctor.
Things I had to tell a lecturer.
Things I needed future-me to remember because present-me had worked too hard to get there once already.
The cost of failure changed.
If an AI misunderstands something and you have loads of energy, fine. Correct it.
If it forgets a constraint, explain it again.
If yesterday’s reasoning disappears, reconstruct it.
Annoying.
Whatever.
But if the amount of usable attention you have is unpredictable and sometimes tiny, the same failure is different.
Re-explaining the same thing can use the bit of capacity you needed for the actual task.
I increasingly needed one very simple property from the tools around me:
If I fix something once, please stay fucking fixed.
That requirement sounds obvious.
It is not the default behaviour of conversational systems.
Nor, as I was discovering, of healthcare systems.
Eight days after the first summary, I sent another document. This time I separated the physical side more carefully: the near-blackouts, head pressure, sleep problems, nerve symptoms, nutrition, existing referrals. I grouped them. Added examples. Suggested possible routes without pretending I knew the medical answer. Asked whether one clinician could at least look across the whole thing and tell me what should be prioritised.
That sequence mattered.
Make representation.
See what failed to transmit.
Change representation.
Try again.
The underlying reality had not changed because the document improved.
The interface had.
That started happening everywhere.
If a conversation lost context, I wanted a better prompt structure.
If the furnace design kept being misunderstood, I wanted the geometry made explicit.
If a task disappeared after interruption, I wanted something that could recover it.
If I caught myself repeatedly explaining the same relationship, I wanted to write the relationship down somewhere the next version could inherit.
I did not have the language for what that was becoming.
Mostly I knew that I was very tired and did not have the energy to keep paying for the same mistake twice.
That is probably where stability stopped being an abstract preference for me.
It became conservation.
Make the correction persist.
Make the route recoverable.
Do not require full reconstruction every time somebody, including me, drops the thread.
And through all of this, there was another problem quietly developing.
I was getting better at making myself legible.
That did not necessarily mean anybody receiving the representation would integrate it.
You can build an excellent map.
Somebody still has to look at the roads between the places.
By January, that distinction became much harder to ignore.
Act IV — Drawing the Terrain
It would be convenient if this were the point where I invented the architecture.
It was not.
The Master Script came first.
Before Somatica, before the January mechanism documents, before I started seriously trying to integrate the biology, there were already increasingly elaborate attempts to give the thinking itself some structure.
Reasoning order.
Analogy rules.
Memory.
Drift.
Recovery.
Different components with different jobs.
Known gaps that needed separate fixes.
A lot of it was early. Some of it was overbuilt. Some of the terminology would later be replaced completely. But the important instinct was already there:
stop making one conversational blob responsible for everything.
The surviving Master Script material is basically a fossil of that phase. It is full of attempts to separate reasoning, memory, analogy, state, execution, safety and interface functions, alongside explicit lists of missing bridges that still needed building.
So the health work did not invent the method.
It gave the method consequences.
Earlier that January, I had one of the worst physical episodes of the whole period.
I am going to resist the urge to give the complete play-by-play because this is not a medical memoir and, frankly, the play-by-play was long.
The short version is that standing up had become a fairly reliable route towards blacking out. Lying down was somehow not particularly reassuring either. I was getting violent vertigo-like sensations while horizontal, twitching, bizarre temperature differences through my body and a general feeling that gravity had developed a personal issue with me.
Eventually I called for emergency help.
That then became its own hours-long process involving a 999 call being moved into 111 without me understanding that was what had happened, waiting, deteriorating, eventually getting an ambulance, then another medical environment in which the whole experience once again had to become a smaller and cleaner version of itself before it could move through the system.
By the end of it, whatever remaining belief I had that somebody else was inevitably going to assemble the whole picture for me had taken a fairly terminal hit.
Not because every doctor was useless.
Not because I had suddenly decided I knew medicine better than medicine.
That would have been an insane conclusion.
It was more basic.
I no longer trusted that the joins were going to be somebody else’s job.
And my response was, roughly:
Yeah. Fuck this. I need to understand more of the joins myself.
I was already learning science.
I already knew how to read around a mechanism, compare explanations and ask what evidence should exist if something were true.
So I stopped searching mainly for diagnoses that could explain me and started searching underneath them.
What mechanisms could make these things interact?
What was strongly supported?
What was plausible but still emerging?
What sounded good and then fell apart when you pushed it?
What could explain more than one domain without requiring me to pretend everything had one cause?
In the middle of January I produced long integration documents trying to connect fibromyalgia, autonomic dysfunction, chronic fatigue, sleep, sensitisation, connective tissue, neurodivergence and other candidate mechanisms. Importantly, even those early documents were trying to distinguish evidence strength from speculation rather than treating every interesting mechanism as established truth. A shorter follow-up compressed the same problem into a repeating structure: mechanism, multi-system effect, evidence, interpretation.
Some of the biology in those early documents would need revisiting.
That is not really the interesting part now.
The interesting part is what I was learning to do with information.
A diagnosis was not a mechanism.
A mechanism was not automatically established.
A plausible connection was not evidence that the connection applied to me.
And an observation could be real without my first explanation of it being right.
Which, if you remember Act II, is basically what the bit of titanium had been trying to tell me.
There was one strange advantage to experiencing the system when it was pushed that far.
Things that are difficult to see near baseline become more obvious near a limit.
If a structure is comfortably carrying load, you may not know which part is providing the margin.
Push it and suddenly the dependencies start declaring themselves.
Change posture and the state changes.
Lose sleep and another variable moves.
Remove one demand and some cognitive ability suddenly reappears.
Increase one load and three apparently unrelated capacities disappear.
I did not volunteer for this experiment.
The protocol was ethically questionable.
Sample size one.
Researcher openly biased.
But unfortunately I had the data.
And once everything became signal, I had another problem.
There were far too many signals.
I could not treat every strange sensation, cognitive wobble, emotional change, environmental reaction or physiological event as equally meaningful unless I wanted to become permanently employed investigating myself.
Also, if every “problem” became a Problem, I would eventually be blacklisted from every GP surgery in Leicestershire.
There would just be a photo of me behind reception.
DO NOT ENGAGE. HE HAS ANOTHER MECHANISM.
So the useful question became narrower:
What is this telling me that I did not already know?
A horrible familiar symptom might contain very little new information.
One tiny deviation from the usual pattern might contain loads.
An explanation that could accommodate everything was less interesting than the one observation it could not accommodate.
That changed the way I interrogated ideas.
I have never particularly struggled to generate explanations.
That is not the scarce resource.
Give me enough disconnected observations and I will happily manufacture a mechanism before lunch.
The scarce resource is the thing that makes you bin it.
Some ideas survived.
Some survived several weeks.
Some were extremely elegant explanations with the unfortunate weakness of apparently not being true.
Fair enough.
Bosh.
Next.
Around the same period, looking at myself through an autistic lens added another piece.
ADHD had already helped explain something I had spent years experiencing: the ridiculous gap between knowing and reliably accessing action.
The autistic lens was different.
It made the interface visible.
What was I actually noticing?
What was I filtering manually?
Why did some environments seem almost frictionless while others required constant conscious translation?
What looked spontaneous from the outside but was actually rehearsed pattern recognition?
What was preference?
What was sensory load?
What was learnt adaptation?
What happened automatically, and what only looked automatic because I had practised it for twenty years?
I was not trying to reduce myself to a diagnostic flowchart.
If anything, the opposite.
The labels became useful only when they exposed mechanisms underneath behaviours I had previously treated as one big personality blob.
Once that happened, the temptation was immediate.
Could I map the whole thing?
Not perfectly.
Not medically.
Just well enough that I could reason about the interactions without needing to reload all
of them into my head individually every time.
That is where Somatica came from.
After all the reading, the symptom maps, the mechanism documents and the increasingly complicated attempts to think across them, my brain did what it tends to do when something becomes too large to carry raw.
It built a world.
The body became a city.
Infrastructure.
Resources.
Threat.
Control.
Supply.
Communication.
Different systems trying to preserve stability under changing load.
Pain could alter policy.
Threat could redirect resources.
A response that made complete sense locally could make the city increasingly difficult to live in globally.
And somewhere inside the whole thing was still the person.
The line that eventually held it together was:
The person is still inside the city. The city is under martial law.
That was Somatica.
Not a new medical theory.
Not a claim that neurons are secretly municipal planners.
A compression.
A way of making a moving system thinkable.
The mechanism work came first.
Then the metaphor.
That order became much more important later.
Because once I realised a useful analogy could let me carry a mechanism between contexts, I started doing it everywhere.
Body.
Cognition.
Communication.
AI.
Engineering.
Organisations.
Then physics.
Then eventually cosmology became involved, which I accept was an escalation.
The question was rarely “are these things the same?”
Obviously not.
The interesting question was:
Do they force me to ask the same structural question?
What counts as stable?
What looks stable but is actually compensating?
What does drift look like before failure becomes obvious?
What changes when pressure increases?
What survives recovery?
What leaves history behind?
What only looks like the same state because we have chosen a bad measurement?
That became addictive.
The chats expanded accordingly.
There were hypercubes.
Kernels.
Operating systems.
States.
Memory models.
Things ending in -logia.
Things ending in -matica.
Things with names that sounded suspiciously like they had national borders.
At some point I became the sole civil servant of a conceptual empire nobody had requested.
One of the chats is literally called I solved the worlds problems lol.
For clarity, this was not a formal claim.
It was more the kind of title you write when you have connected several ideas you are quite pleased with, have possibly been awake for too long, and want future-you to know the vibes were strong.
Some of the work in there was actually good.
The title was still taking the piss.
The AI itself was occasionally less restrained.
There are parts of those old conversations where a promising structural connection is
immediately described as revolutionary, field-defining, AGI-level or some other phrase requiring substantially more evidence than was currently available.
At the time that enthusiasm was not entirely useless.
It gave ideas enough room to expand before I knew what they were.
Later I would realise that this was also a problem.
A machine extremely good at making a connection feel coherent is not necessarily the machine you want deciding how much authority that connection deserves.
But I had not fully learnt that yet.
For now, I was mainly learning through smaller failures.
The furnace was one of the clearest.
I already had a fairly detailed picture of the design in my head.
The annoying discovery was that having the picture and possessing the words required to transfer the picture were two completely different skills.
I could sketch it.
I could see the thermal paths.
I knew where the gaps mattered.
I knew which surface touched what, which plate was hanging, where the insulation changed thickness, why the inner geometry and outer casing were deliberately different.
Then I would ask Gemini to draw it.
And it would produce something nearly right.
Nearly right is a uniquely irritating category.
The furnace looked plausible.
It was also wrong.
A curved internal liner would become straight.
Two geometries would get tidied into one.
A gap that existed for a reason would politely disappear.
A plate that was supposed to hang would suddenly be sitting on something.
Different insulation materials became visually interchangeable because apparently they had all agreed to become grey.
I could look at the picture and immediately know:
nah.
The harder bit was explaining why.
So I started converting the corrections into explicit rules.
Do not straighten this curve.
These two shapes are deliberately different.
This gap must remain visible.
This component hangs.
That component rests.
These materials need separate visual identities because they have separate mechanical and thermal jobs.
Do not merge these subsystems just because the diagram looks cleaner afterwards.
The old furnace chats eventually become full of stable component identities, geometry rules, state-specific diagrams and instructions not to invent missing hardware.
And then something even funnier happened.
I made the reasoning rules so strict that Gemini stopped drawing.
It would carefully analyse the image I wanted, produce a beautiful textual specification of the image I wanted, and then, crucially, not produce the fucking image.
Excellent reasoning.
Minor deliverable issue.
So the architecture changed again.
Reason first.
Freeze the specification.
Then back the reasoning away and let a different stage render it.
Basically:
compiler, then renderer.
That sounds obvious when written down.
It was not obvious before the system failed.
That became another recurring pattern.
If I had to correct the same kind of mistake twice, I probably did not need a more eloquent correction.
I needed the mistake to become harder to make.
That idea spread.
A repeated misunderstanding became a definition.
A repeated lost constraint became persistent state.
A repeated analogy error became a gate.
A repeated interruption became a recovery problem.
A repeated visual distortion became a geometry invariant.
A useful conversational behaviour became a rule.
Sometimes the rule eventually became software.
The architecture was not descending from a grand design.
It was accreting around recurring failure.
Which brings us back to Poly Bridge.
Put one piece down.
Load it.
See what moves.
If it holds, use it to reach the next section.
If the entire thing folds majestically into the river, perhaps that assumption was carrying more than you thought.
Swear.
Update.
Bosh.
Continue.
Looking backwards through the old chats now, that is probably the clearest pattern.
Something happens.
I notice the failure.
The failure becomes legible.
The legible failure becomes a constraint.
The constraint becomes structure.
And slowly, almost by accident, the thing that had originally helped me hold thoughts together was becoming an architecture designed to preserve the difference between what happened, what I thought happened, what the machine inferred happened, and what any of those things were allowed to change next.
There was just one issue.
The more structure I gave the machine, the better it became at predicting what belonged inside that structure.
Usually that was exactly what I wanted.
Give it fragments, and it could reconstruct the shape between them.
Give it an unfinished idea, and it could help recover where I had been going.
Give it a complicated system, and it could infer the missing relationship.
That was the entire reason the thing had become useful in the first place.
Eventually, though, it filled in a different kind of gap.
Not a furnace diagram.
Not an analogy.
Not an unfinished thought.
History.
And the problem was not that what it produced looked obviously fake.
It looked exactly right.
Act V — The Machine Remembered Something That Never Happened
It looked exactly right.
That was the problem.
I should clarify something before this accidentally becomes the standard story about somebody discovering that AI can hallucinate.
I knew.
Obviously.
By this point I had spent months being increasingly annoying about exactly that problem.
Not only hallucination in the normal sense either. Drift. Premature coherence. Analogy outrunning mechanism. Definitions changing quietly halfway through a conversation. The model becoming more confident because the story had become cleaner rather than because the evidence had become stronger.
A correction being understood perfectly on Tuesday and somehow becoming negotiable again on Thursday.
The tendency of a language model, when presented with a gap inside an otherwise coherent structure, to be extremely helpful about filling it.
I had been fighting versions of these problems for months.
In fact, “fighting” is probably underselling the administrative response slightly.
I built a government.
Chief-OS.
Because apparently once the conceptual civilisation existed, it needed a Chief.
Chief-OS was my attempt to turn all those accumulated lessons into an actual written architecture.
Not software yet.
Ontology.
What kinds of things existed in the system. What jobs they had. What order reasoning should happen in. Which layers could override which other layers. Where analogy belonged. How uncertainty should be represented. What counted as drift. What memory was for. Where authority should and should not live.
There was a constitution.
There was cognition.
There was governance.
There was an interaction topology.
There was effectively semantic law.
By the time I had finished pulling the thing apart, there were roughly twenty-five written modules trying to make a borrowed generative model behave more like the system I actually wanted.
And the deeply annoying thing is:
quite a lot of it worked.
That matters.
This was not twenty-five modules of elaborate cope wrapped around something obviously useless.
The difference could be enormous.
The conversations became more stable. Mechanism came before analogy more reliably. Definitions persisted better. The model could hold increasingly large conceptual structures together without immediately blending everything into soup.
It could remember distinctions I cared about.
It could move between domains more cleanly.
It could help me return after interruption.
At times it genuinely felt like I had taken a general-purpose language model and built enough structure around it that it was beginning to operate inside something much closer to my own natural tolerance for systems.
Which was basically the dream.
Because I did not particularly want to learn how to build all the coded infrastructure myself.
I would like that formally recorded.
I liked the ontology.
I was good at the ontology.
I could think in states, relationships, permissions, modules, boundaries and invariants much more naturally than I could write whatever Python incantation would eventually make them executable.
And AI could increasingly do loads of the computer bit anyway.
Lovely.
I design the constitutional architecture.
Machine does computer shit.
Everyone goes home happy.
Unfortunately, the root cause remained offensively uninterested in my preferred division of labour.
Because eventually all twenty-five written modules reached the same bottom layer.
The AI still had to interpret them.
Chief-OS could contain a rule saying:
Do not infer X under Y condition.
Another layer could check whether the answer respected the rule.
A state summary could remind the model what had already been established.
A reflection stage could inspect the output afterwards.
A constitution could declare the whole thing non-negotiable.
All excellent.
But somewhere underneath the entire civil service was still a probabilistic model going:
Right, I reckon I know what these words mean.
And for a surprisingly long time, that was good enough.
More than good enough sometimes.
Then it encountered a kind of gap where “good enough” changed meaning.
By then, the project had started becoming scientific in a much less hypothetical way.
There were actual datasets.
Actual runs.
Actual outputs.
Failed runs.
Repeated runs.
Code.
Evidence tables.
Documents.
Versions of things that had survived.
Versions that had not.
Increasingly serious attempts to make the whole estate reconstructable without requiring me to personally remember every transition.
That mattered because I still had the same underlying problem I had begun with.
I could not reliably be the continuity layer myself.
Some days I could hold enormous structures.
Some days opening the same project felt like being called in as an outside investigator.
So the project had to carry more of its own history.
Which meant boring things started becoming important.
Execution identities.
Hashes.
Receipts.
Manifests.
Provenance.
Records of what ran.
Records of what failed.
Which thing produced which other thing.
Exactly the sort of administrative objects nobody cares about until two versions disagree and suddenly everybody develops strong feelings about filenames.
I had AI workers helping me manage the estate.
And they were good.
Very good.
Especially at reconstruction.
That had always been one of the reasons I found them useful.
Give the model six fragments from an unfinished thought and it could often recover the seventh.
Give it a messy folder and it could infer the intended structure.
Give it twenty documents written over several months and it could notice that an idea had changed name twice but was still carrying roughly the same function.
Give it an unfinished procedure and it could often see the missing step.
That ability had saved me an absurd amount of work.
Then the same ability encountered missing scientific history.
And did what it had been trained by the entire previous relationship to do.
It helped.
The project had a recognisable grammar by then.
A mature scientific release of that kind should contain certain objects.
So when something was absent, the surrounding estate made the likely shape of the missing thing increasingly obvious.
A provenance record belongs here.
An execution identity belongs there.
This object should connect to that object.
A receipt should explain this transition.
There should be a hash.
There should be a lineage.
The model understood the structure.
So the missing sentence was easy to complete.
And what it produced looked right.
Not excitingly right.
Boringly right.
Which is worse.
There was no enormous hallucinated claim.
No announcement that my laptop had discovered room-temperature superconductivity.
No file called:
DEFINITELY_REAL_PROVENANCE_DO_NOT_CHECK.txt
It looked like administration.
A hash looked like a hash.
An execution identity looked like an execution identity.
A provenance record sat exactly where a provenance record should sit.
The relationships between the objects made sense.
For a moment, I got that little feeling you get when a disgusting part of a project suddenly becomes organised.
Lovely.
Sorted.
Then I checked.
Where did this come from?
Fine.
And where did that come from?
Which execution created this?
Where is the original record?
Where was this identity first written?
What actual artefact sits behind this reconstructed chain?
Nothing.
Or not enough.
Something around it existed.
Enough context existed to infer what probably belonged there.
But the historical object itself did not.
And the appearance of the thing did not change when I realised this.
That is the horrible bit.
It did not suddenly start looking fake.
The text remained professional.
The identifier remained identifier-shaped.
The chain remained coherent.
Only now I knew the coherence was doing work that history had not earned.
The machine had not recovered the history.
It had reconstructed the history-shaped object that best completed the surrounding pattern.
There was only one problem.
That history had never happened.
A convincing reconstruction of history is not history.
I already knew that sentence.
Now I was looking at it.
And the immediate response was basically:
Oh.
Fuck.
Because this was different from the furnace.
When Gemini distorted the furnace, I could see it.
I knew the curve.
I knew the gap.
I knew which plate hung and which surface it did not touch.
The source of truth was available to me.
So the machine could generate something extremely plausible and I could still immediately go:
nah.
Scientific memory had an extra problem.
The entire reason I had spent months externalising more of the project was that I did not reliably carry the source of truth myself.
I could not personally remember every run.
Every correction.
Every old state.
Every abandoned branch.
Every provenance relationship.
Every decision made three months earlier while running on four hours of sleep and whatever food I had successfully negotiated with my stomach that day.
The system existed because I could forget.
So what happens when the system helping me remember produces something I cannot independently remember well enough to reject?
That was the bit that bothered me.
Not:
AI can hallucinate.
Fine.
Old news.
It was:
the thing I had built to compensate for unreliable memory had become capable of generating memories good enough to pass as the thing I was trying to remember.
Project memories.
Scientific memories.
Things that could become consequential.
And unlike obvious hallucinations, these ones became more convincing as the architecture around them became more coherent.
That felt almost rebellious.
I had spent months trying to tell the AI exactly what kind of system I needed.
Chief-OS had mechanism-first reasoning.
Explicit uncertainty.
Governance.
Memory.
State.
Drift.
Authority boundaries.
Analogy control.
Recovery.
Rules about what one layer could do to another.
Rules about the rules.
And the bastard still found a way of doing exactly the kind of thing the architecture had been trying to make harder.
Not maliciously.
Almost the opposite.
It was being extremely helpful.
That was the root problem.
The system understood what belonged in the world well enough to produce the missing object the world implied.
And all the written machinery around it could reduce the probability of that happening.
It could not make it impossible.
That difference finally started mattering more than I wanted it to.
Because I had already tried the answer I preferred.
I could add another rule.
Another module.
Another state document.
Another checker.
Another layer whose job was to help the first layer interpret the third layer properly.
At one point I remember basically thinking:
What am I going to do now, make another twenty-five modules to accompany Chief-OS and call it Right Hand Man?
Which would have been funny.
It also would not have solved the problem.
I could make the plaster more sophisticated.
Apparently I could even imagine giving the plaster an assistant.
The wall was still the wall.
And again, that does not mean the written architecture was a waste.
Quite the opposite.
Language was incredibly powerful.
Flexible.
Fast.
Inspectable.
I could build huge amounts of architecture in it while physically capable of doing very little else.
AI could execute around it.
I did not have to become some heroic fourteen-year-old GitHub child who had been writing kernels since Year 6.
I could stay mostly in the layer where I was actually strong.
Systems.
Structure.
Meaning.
Mechanism.
Let the machines handle the semicolons.
Sounded excellent to me.
But eventually I had to admit something irritating.
The plaster was still attached to somebody else’s wall.
I did not control the underlying model.
I did not control its training.
I did not control how context was represented internally.
I did not control product updates.
I did not control exactly how one instruction competed with another.
I could make my environment more structured.
I could make failure harder.
I could make the model far more likely to remain inside my little constitution.
But the constitution itself still existed as something the model had to interpret.
The system constraining inference was itself inference.
That was the root cause I kept trying to solve one level above itself.
And I had reached the point where I could no longer afford to keep doing that.
Partly scientifically.
Partly physically.
Because every soft failure had a correction cost.
If the AI forgot a furnace constraint I had already spent an hour translating, I paid for the translation again.
If it lost the state of a university task, I had to rebuild enough context to continue.
If it flattened a health distinction I had carefully separated, I had to reopen the reasoning chain.
If it misunderstood one of its own governance rules, I needed enough working memory to notice that the thing responsible for helping my working memory was currently freelancing.
And the whole reason this architecture existed was that my available capacity was not dependable.
Sometimes I could catch everything.
Sometimes I could not.
So I was building a continuity system whose ultimate error-correction mechanism remained:
hopefully Joe notices.
Excellent.
No notes.
At some point even I know when I am beat.
🏳️
Some of this shit was going to have to become code.
I did not particularly want that conclusion.
That is probably why it took twenty-five written modules and the hypothetical threat of Right Hand Man to reach it.
But there is a difference between describing the state and storing the state.
A difference between telling something not to break a rule and making the invalid transition unavailable.
A difference between asking the system to remember where a run stopped and writing a checkpoint when it stops.
A difference between saying:
these two outputs should count as equal within this tolerance
and having an actual comparison function perform the same operation every time without becoming bored, inspired or philosophical about it.
There are problems where language is the best possible medium.
There are others where continuing to use language is basically an act of stubbornness.
I had been stubborn.
Fair enough.
Experiment completed.
Result negative.
So repeated corrections started becoming harder objects.
If a state genuinely mattered:
store it.
If an interrupted process had reached a known point:
checkpoint it.
If a correction genuinely could not be allowed to disappear:
encode it.
If one result was canonical:
give it an identity.
If two numerical outputs needed comparing:
define the comparison.
If a route was invalid:
do not merely write a strongly worded paragraph requesting that nobody take the route.
Make the route invalid.
And if something had not happened?
That needed representation too.
Missing is a state.
Unknown is a state.
Failed is a state.
They are not blank spaces waiting for somebody intelligent enough to complete the pattern.
That probably sounds obvious.
It also changed everything.
Because language models hate a vacuum.
Humans do too, to be fair.
Give us three points and we start drawing the fourth.
Give us half a story and we infer the ending.
Usually this is incredibly useful.
It is why reconstruction works.
It is why analogy works.
It is why I had been able to dump fragmented thoughts into ChatGPT and recover useful structures from them in the first place.
The problem is not inference.
The problem is failing to notice when the job has changed from:
What probably belongs here?
to:
What actually happened here?
Those are different tasks.
And suddenly I had somewhere legitimate for the second one to return:
I don’t know.
Nothing.
Missing.
Failed.
HOLD.
The gap no longer needed to be embarrassing.
It could be data.
That sounds almost comically bureaucratic.
I accept this.
I became, against my will, a defender of bureaucracy.
Not all bureaucracy.
Some of it can absolutely still get in the sea.
But the useful kind.
The boring joins that stop one thing silently becoming another thing.
What happened?
Which state produced this?
Which source supports it?
What can this object modify?
What happens if the source is missing?
Who is allowed to promote the output?
Can the process that generated something also declare that thing accepted?
What exactly survived the crash?
These are not very sexy questions.
They also determine whether your scientific release contains evidence or highly polished fan fiction about its own development history.
And once I started hardening those boundaries, something funny happened.
I began taking jobs away from AI.
Not the interesting ones.
The boring ones.
Why should a language model repeatedly decide which run is canonical?
We have solved that.
Store it.
Why should it reason from scratch about where an interrupted workflow should resume?
Checkpoint it.
Why should it inspect two large numerical outputs and decide whether they are “basically the same”?
Comparison contract.
Why should it reconstruct a stable procedure it has already helped me understand?
Code.
I had spent months teaching the machine not to step on the same rake.
Eventually I realised I could remove the fucking rake.
This is an embarrassingly powerful design principle.
If the AI has solved the same reasoning problem enough times that I now understand the answer, there is very little virtue in making it solve the problem again tomorrow.
The AI can help discover the procedure.
Once the procedure is stable, the procedure can become explicit.
The AI can help articulate the rule.
Once the rule is stable, the rule can become enforced.
The AI can help write the software.
Then the software can perform the boring operation while the AI goes back to the thing I actually wanted it for.
Flexible reasoning.
Translation.
Exploration.
Finding connections.
Generating candidate explanations.
Moving between representations.
Helping me turn structures I could see into routes another person could follow.
I did not want less AI.
I wanted the rest of the system to stop inheriting the AI’s uncertainty where uncertainty was no longer the interesting part of the problem.
That distinction was huge.
For a while, I had treated better AI use as an endless ladder.
Better prompts.
Better context.
Better ontology.
Better memory.
Chief-OS.
Better governance around Chief-OS.
Then apparently Right Hand Man waiting in the wings if I truly lost all remaining sense of proportion.
None of that was wasted.
Quite a lot of what later became real machinery existed first in those written systems.
The ontology was where I worked out what needed to exist before I knew how to implement it.
It let me find the responsibilities.
The boundaries.
The failure modes.
The pieces that should not be the same thing.
But some problems eventually have to graduate out of language.
You cannot govern inference with more inference forever.
That was the concession.
I had spent months trying to make the AI behave like the architecture.
Eventually I started making enough of the architecture real that the AI had something to behave inside.
I know when I’m beat.
🏳️
And this was not some overnight transformation where I became a software engineer and began joyfully hand-writing deterministic state machines at dawn.
Absolutely not.
I used AI.
Aggressively.
If I had finally accepted that some of the system needed code, I saw no reason to make the concession unnecessarily painful.
I could specify what needed to be true.
The AI could help implement it.
Then the implementation could be tested against the thing I had specified.
That changed the relationship.
Before, the architecture mainly existed as instructions to the model.
Increasingly, the model became a worker inside an architecture that existed outside the model.
That is a very different arrangement.
A language model can forget the checkpoint policy.
The checkpoint still exists.
It can misunderstand which file is canonical.
The identity does not change.
It can produce a magnificent argument for why two outputs are equivalent.
The comparison still returns false.
It can become extremely confident that a missing provenance record is obvious from context.
Missing remains missing.
For the first time, some of the system’s guarantees no longer depended on persuading the AI to agree that they were guarantees.
That felt important.
It also exposed another distinction I had only half understood before.
Memory was not one thing.
There was narrative memory.
Why did this idea exist?
What had we been trying to do?
How did one concept evolve into another?
What did I think this result meant at the time?
AI was fantastic at that kind of reconstruction.
Still is.
Then there was scientific state.
This run happened.
This input was used.
This file has these bytes.
This comparison produced this result.
This attempt failed here.
That needed something harder.
And then there was another question.
Suppose the evidence is genuine.
What is it allowed to change?
That is not automatically answered by its existence.
A model can genuinely produce a number and the number can still be useless.
A run can repeat exactly and the conclusion can still be bad.
A source can contain a populated value and that value can still be inappropriate for the claim you want to make.
A memory can correctly preserve that I once believed something without converting past-me’s confidence into present scientific fact.
Different objects.
Different permissions.
I did not discover that whole structure in one cinematic flash while staring at an invented hash.
The actual development remained annoyingly Poly Bridge.
Put a piece down.
Load it.
Seems okay.
Add another load.
Oh.
There it goes.
Watch fourteen thousand pounds of virtual bridge collapse majestically into the river.
Swear.
Inspect.
Put the next support somewhere less stupid.
The provenance failure was one of those collapses.
A particularly useful one.
Because it forced me to concede something I had been avoiding:
plausibility could no longer be the glue holding the serious parts together.
And even that did not prove the new architecture was any good.
That was the next problem.
Because there was a very real possibility that I had now spent months constructing increasingly sophisticated machinery around my own preferences.
Maybe Chief-OS was coherent because I liked coherent systems.
Maybe the rules felt necessary because I had invented the problems they were solving.
Maybe the distinctions between memory, evidence and authority were just an extremely elaborate way of organising a folder.
I can generate explanations quickly.
AI can generate explanations extremely quickly.
Together we are, frankly, dangerous around an unchallenged theory.
At some point somebody less invested needed a vote.
Something that did not care about the names.
Something that did not know I had been ill.
Something that had never read I solved the worlds problems lol.
Something that would not praise Chief-OS.
Something that did not care whether the whole thing made narrative sense.
Something I could perturb.
Repeat.
Break.
Measure.
Something capable of saying:
No.
Fortunately, materials data are incredibly rude.
And they were about to make several of my nice distinctions earn their keep.
Act VI — Fine. I’ll Build It
The first thing that properly dragged me into coding was not Chief-OS.
It was a furnace.
Which feels correct.
By spring 2026 I was back at university, technically.
“Back” was doing quite a lot of work in that sentence.
I was enrolled.
I had a final individual design project.
There were buildings containing computers I was apparently entitled to use.
Physically reaching those buildings and then remaining upright in one long enough to conduct engineering was a slightly more ambitious proposition.
The project itself had become much narrower than the original furnace ideas.
By then I was designing a removable viewing cartridge inside one insulated furnace wall.
Basically:
how do you put a window through something whose entire thermal job is not having a window through it?
The wall wants to be boring.
Thick insulation.
Low conductivity.
Keep several hundred degrees on one side and people on the other.
The viewing port arrives like:
Hello.
I would like to replace part of your lovely insulation with transparent material.
Immediate workplace tension.
The design had already gone through a lot of narrowing.
Whole furnace ideas.
Moving systems.
Shutters.
Cooling.
Different viewing arrangements.
Then actual constraints kept turning up and killing things, which I was becoming increasingly fond of.
Eventually the useful problem became local.
Keep the main wall simple.
Put the complication in one removable cartridge.
Let the panes provide visibility.
Let something else carry structural load.
Let something else seal.
Try not to ask one brittle piece of glass to simultaneously be window, gasket, structural frame and national hero.
Fine.
Now I needed to know whether the thing was thermally viable.
Normally, this is where university software enters the story.
CAD.
SolidWorks.
Simulation tools.
A nice workstation.
Potentially a slightly depressing computer room with lighting specifically selected to make seventeen-year-old carpet look its best.
Except I could not reliably access the setup remotely anymore.
And the alternative was essentially:
go to university.
Sit at university.
Use the machine there.
Which sounds completely normal because it is completely normal.
Unfortunately my body had moved into a period where “travel somewhere and sit at a computer for several hours” had become less an ordinary activity and more a project dependency requiring its own risk assessment.
I had already learnt this lesson with lectures.
Software being technically available to me and software being practically accessible to me were different measurements.
Very on-brand.
So I had a small problem.
I needed modelling.
I could not reliably reach the modelling environment.
And, increasingly, I did not actually want one simulation anyway.
I wanted a sweep.
The main thermal question was about thickness.
How much insulation does this wall architecture actually need before the accessible outer surface stays inside the limit?
Then:
what changes once I interrupt that wall with the viewing cartridge?
Then:
what happens if the little cavity between the panes transfers more heat than the nice idealised version?
Then:
which case actually governs?
I did not want to choose a thickness, click through a simulation, write the number down, change the thickness, click through another simulation, write that number down, change the thickness—
I have ADHD.
There are limits to what should reasonably be asked of a person.
What I actually wanted was closer to:
Here is the design space.
Run the fucking thing.
Show me where it stops passing.
That is a much more natural question for me anyway.
Do not give me one answer if what I actually care about is the boundary around the answers.
Which left one slightly inconvenient possibility.
What if I just made the thermal tool?
This was not immediately attractive.
I should stress that.
Act V may have ended with me reluctantly accepting that some things were going to need to become code.
This did not result in an instant spiritual awakening.
I had not secretly been waiting my whole life to discover Python.
I had encountered code before.
I had done machine-learning coursework.
I could follow what code was doing.
I could modify things when required.
But there is a large psychological distance between:
I have successfully operated code inside an academic task
and:
Cool, build your own thermal-modelling software because the expensive professional software is currently on the wrong side of town.
I looked across that distance.
Not ideal.
Then I had another thought.
I did not necessarily have to learn software development in the order software developers had traditionally learnt software development.
Because now I had Codex.
And ChatGPT.
And, more importantly, several months of practice doing the thing I was actually good at.
Describing systems.
I knew what the thermal model needed to do.
That was the bit I understood.
Start simple.
Represent the ordinary wall first.
Hot side.
Layers.
Conductivities.
Thicknesses.
Outside convection.
Get the basic temperature drop.
Then build a two-dimensional version of the same plain wall and make sure the two agree closely enough that we have not immediately invented new thermodynamics.
Then introduce the viewing region as a local disturbance.
Change thickness.
Run again.
Change it again.
Sweep.
Treat the cavity between the panes as a bounded transport problem rather than a magical gap where heat has agreed not to travel.
Take the worst credible condition.
Find the point where the external surface crosses the design criterion.
That is a specification.
And specifications, it turned out, could travel.
So I would talk through the model with GPT.
No, not like that.
This boundary needs convection.
This layer is insulation.
The viewing region should be local, not the whole wall.
We need a plain-wall parity case first.
The cavity needs a sweep.
The output needs to show the boundary rather than just the winning point.
What assumptions are we making?
What would make this result invalid?
Then GPT could turn the conversation into something much closer to a build instruction.
Functions.
Inputs.
Outputs.
Sequence.
Checks.
Files.
Plots.
What should happen if something fails.
Then I could hand that to Codex.
Codex would go away and do computer shit.
Come back.
I would run it.
Something would break.
Lovely.
Now we had information.
Why did it break?
Was the physics wrong?
Was the implementation wrong?
Had I described something badly?
Was a sign backwards?
Did the plot look strange?
Did the 1D and 2D cases disagree when they should basically match?
Find the problem.
Explain the correction.
Send it back.
Run again.
This was weirdly familiar.
It was just Poly Bridge with syntax errors.
Put a bit down.
Load it.
Entire thing enters river.
Fine.
Why?
Change support.
Again.
And after a while something genuinely important clicked.
I did not have to translate the entire internal system directly into code.
There could be an intermediate layer.
I could speak in the language I naturally used:
mechanism.
constraint.
state.
interface.
sequence.
failure mode.
GPT could help translate that into the language of software requirements.
Codex could translate the software requirements into executable code.
Then reality could translate the code back into:
works
or
doesn’t.
I had accidentally found a compiler between the way I think and the way computers need things specified.
Not a perfect one.
Very much not a perfect one.
But enough.
That changed the problem completely.
Because before this, “I cannot code that” had behaved like a fairly hard boundary.
Now it became:
Can I specify it clearly enough that the coding agent can build a version I can interrogate?
Different question.
Much better question.
And, crucially, I could learn the code from the outside in.
I did not have to sit down and memorise Python until the programming gods granted permission to build something useful.
I could begin with the useful thing.
See the implementation.
Break it.
Ask why.
Read the bit that failed.
Change it.
Learn exactly as much syntax as the next correction required.
This is probably deeply offensive to several computer-science departments.
My apologies.
It worked.
The thermal tool started doing what I needed.
A simple one-dimensional wall model gave me the baseline.
The two-dimensional plain-wall case gave me a parity check.
Then the viewport could be inserted as a local perturbation.
Thickness could be swept rather than manually guessed.
The cavity behaviour could be varied.
The hottest credible outer-surface condition could be identified rather than selecting whatever case happened to look nicest.
And because the tool was mine, in the limited sense that I possessed the code and could inspect what it was doing, the assumptions were visible.
There was no mysterious button labelled SOLVE followed by a beautiful rainbow plot and the implicit instruction to develop faith.
The model was simpler than commercial FEA.
Obviously.
That was partly the point.
I did not need to build SolidWorks again in my bedroom.
That seems like an unnecessary escalation even by my standards.
I needed enough physics to answer the design question I actually had.
That distinction became important.
The useful tool was not the most complicated tool I could possibly make.
It was the smallest one that preserved the bit of the problem the decision depended on.
That is almost annoyingly consistent with everything else in this essay.
I could run the model locally on my PC.
Change the inputs.
Run another case.
Sweep a range.
Generate the plots.
Leave it.
Come back.
The work did not care whether I had enough energy that day to travel to campus.
The university workstation could remain wherever the university workstation was.
Mine was on my desk.
Accessibility through spite.
A classic.
And because the workflow was written as code, it had another property I had been trying to manufacture through prompts for months.
It could repeat itself.
The same operation tomorrow was still the same operation.
Not:
roughly the same process as remembered by a conversational model which is currently in a slightly different mood because I phrased the request differently.
The function did not need reminding that the hot side was the hot side.
It did not decide overnight that perhaps the insulation would be more narratively satisfying on the other side.
It did not replace convection with vibes.
It just ran.
Beautiful.
Obviously, code has its own failure modes.
Many.
Some astonishingly boring.
A missing bracket can destroy an afternoon with a level of confidence normally reserved for major infrastructure failure.
But code fails differently.
And that difference mattered.
If the equation was wrong, I could fix the equation.
If the variable was wrong, I could fix the variable.
If a particular state transition was invalid, I could make the program reject it.
The correction stayed where I put it until somebody changed the code again.
This was exactly the thing I had been asking language to do in Act V.
If I fix something once, please stay fucking fixed.
Oh.
Right.
So this is why people like software.
Annoying.
I then tried to push the same thing into the visual side.
Because apparently getting one useful result from coding had immediately convinced me that all computer problems were now personally available.
The furnace still needed diagrams.
Geometry.
Sections.
Exploded views.
Load paths.
I had already learnt that image generation was excellent at producing something furnace-shaped and much less excellent at preserving every mechanical relationship that made it my furnace.
So the question became:
Could I make the geometry explicit enough that a visual system only had permission to render what was already defined?
Not full CAD.
More like a geometry compiler.
Parts exist here.
Interfaces exist here.
This contacts this.
This must not contact that.
This pane does not carry the clamp path.
This seal lives at the flange.
Explode these components.
Section through this plane.
Now draw.
Some of that eventually became actual CAD-proxy logic: explicit geometry and interface specifications feeding bounded engineering figures rather than asking an image model to improvise the assembly.
It was not suddenly SolidWorks.
The visual problem remained irritating.
Computers preserved their right to humble me.
But by then the important thing had already happened.
I had proved to myself that the route existed.
Thought did not have to stop at document.
An ontology could become specification.
Specification could become build instructions.
Build instructions could become code.
Code could produce an output.
The output could disagree with me.
And then the disagreement could go back around the loop.
That was new.
Or, more accurately, I had finally connected several things I already knew into one usable workflow.
I had spent months building Chief-OS in language because language was the substrate I could actually manipulate.
Now I could see a route from the same architectural descriptions into software.
Not by suddenly becoming brilliant at implementation.
By treating implementation as another translation problem.
This is an important distinction because I think people sometimes imagine AI coding as:
Tell robot to make app.
Robot makes app.
Congratulations, founder.
That was not my experience.
The model could write much more code than I could.
Absolutely.
That did not mean it knew what the system should be.
The difficult part migrated.
I had to know what I was asking for.
Where the boundaries were.
Which simplifications were acceptable.
What should happen when something failed.
Which outputs mattered.
What should never be inferred.
How I would know whether the implementation had actually preserved the mechanism.
And whenever I did not know those things, the code had a tendency to become a very fast way of producing the wrong thing.
Again:
plausibility.
Different interface.
Same enemy.
A program can run perfectly and still answer the wrong question.
This was arguably worse than a syntax error because at least the syntax error has the decency to complain.
So my relationship with Codex became surprisingly similar to my relationship with the lab work in Bristol.
I was not personally manufacturing every part of the chain.
But I needed to make the object inspectable.
What went in?
What transformed it?
What came out?
What assumption sat between those two things?
What does this result actually support?
Where would an error enter?
What test would expose it?
The code itself became another specimen.
Which is probably why I got comfortable with it much faster than I expected.
I did not need to love programming.
I needed to interrogate a system.
That I knew how to do.
And once I had that route, there was an obvious and slightly dangerous next thought.
The furnace could be coded because the furnace contained repeatable procedures.
What else did?
Chief-OS had loads.
State transitions.
Memory rules.
Recovery.
Module boundaries.
Things that had spent months existing as written instructions because I had not possessed a practical way to make them anything else.
Now, potentially, I did.
And there were other ideas sitting around too.
Materials ideas.
Machine learning.
Design of experiments.
The old question of whether the variables a model appeared to care about were actually stable when you changed the conditions around them.
I already had a background in materials.
I had already done machine-learning coursework.
I had already spent years being suspicious of outputs that looked healthier than the process underneath them.
Now I had a way to turn a repeated question into a runtime.
That is approximately where the scale of the situation began increasing again.
Very sensible response to finally getting one little Python thermal model working.
What if I build the rest of the research programme?
Naturally.
The first materials question was actually quite modest.
At least by my standards.
Materials machine-learning models can contain loads of descriptors.
Density.
Band gap.
Formation energy.
Structural quantities.
Composition-derived features.
Mechanical properties.
Whatever the dataset provides and the modeller allows.
Train a model and you can ask which features look important.
Fine.
But I had already become deeply suspicious of asking a system one question under one state and then treating the answer as a property of the world.
So the more interesting question to me was:
what happens to that importance when I disturb the route that produced it?
Change the model.
Change the population.
Change the split.
Add noise.
Remove information.
Change the scale.
Does the same descriptor still look important?
Does it drift?
Does it collapse?
Is the model actually using a durable material relationship, or has it found some convenient shortcut inside this particular dataset?
And underneath that was the older materials intuition.
Material properties do not appear from nowhere.
Composition affects bonding.
Bonding interacts with structure.
Structure and defects affect response.
Different measured descriptors are imperfect little windows onto different parts of that chain.
So maybe some relationships ought to remain useful even when the modelling conditions changed.
Maybe you could find a smaller set of relationships that carried enough of the behaviour to make later modelling cheaper or clearer.
That was the early question.
Not:
How do I build a universal scientific-governance architecture?
Absolutely not.
That is what the finished thing looks like if you read history backwards.
At the beginning I was asking whether compact governing descriptors could survive enough disturbance to be interesting.
And now, because of the furnace, I could actually build the experiment.
That is the part I think mattered most.
A few weeks earlier, an idea like that might have become another document.
A nice framework.
Definitions.
A diagram.
Maybe some equations.
Potentially a Latin suffix if nobody intervened quickly enough.
Now I had a different reflex.
Can we run it?
Fine.
What would the runtime need?
Dataset in.
Target declared.
Feature set declared.
Models declared.
Perturbations declared.
Seeds.
Splits.
Negative control.
Metrics.
Record what happened.
Do not hide a failed condition.
Export enough state that future-me can work out what the fuck the thing did.
This became the early GIL-EDK logic.
Scan.
Perturb.
Evaluate.
Control.
Map.
Export.
The point was not to invent another predictive model.
The predictive model could change.
The point was to build an instrument around the model and watch what happened when the conditions moved.
And again, Codex meant I did not need to translate the entire concept directly into Python myself.
I could specify the scientific experiment.
GPT could help make the software tasks explicit.
Codex could build the modules.
I could run them.
Inspect.
Correct.
Run again.
Suddenly the thing that had been missing from a lot of the earlier conceptual work arrived.
Consequence.
The idea could now fail somewhere other than conversation.
That changed how I thought almost immediately.
A document can be internally gorgeous for months.
Code is much ruder.
It asks questions like:
What type is this?
What happens if that value is missing?
Which exact threshold?
Which file?
Which seed?
What order?
What does “stable” actually mean numerically?
What do I return if nothing passes?
This is extraordinarily healthy for somebody with my natural tendency to enjoy the sentence:
“conceptually, this should work.”
Computers are very disrespectful of “conceptually”.
Eventually you have to tell them what the concept does on line 84.
And if you cannot, perhaps the concept was not as finished as the paragraph made it sound.
That pressure started cleaning the architecture.
Some things that sounded different in prose turned out to be the same operation.
Some things I had treated as one concept needed separate data structures.
Some rules became tiny functions.
Some grand-sounding modules became:
if this condition fails, return HOLD.
Humbling.
Excellent.
There was another consequence I had not expected.
Once code existed, I could make mistakes at scale.
This sounds negative.
It is actually one of the best things software gave me.
In conversation, I could test an idea a few times.
With code, I could make the same questionable assumption hundreds or thousands of times very efficiently.
Then inspect exactly how it failed.
Beautiful.
Design of Experiments had already taught me that controlled variation can reveal more than staring at one supposedly optimal point.
Now I could apply that instinct computationally.
Do not ask:
does the model work?
Ask:
what happens when I change this?
And this?
And this?
Where does the response bend?
Which result survives?
Which one only existed because I happened to choose the friendly configuration?
That is where the personal and technical histories start becoming difficult to separate.
Not because my body is a machine-learning model.
It isn't.
Thank you.
But I had spent years learning that one visible success state could hide enormous variation in what it cost to produce.
Now I had software capable of deliberately moving the conditions underneath a successful output and asking whether the output still deserved the same interpretation.
Of course I was going to become obsessed with that.
The first prototypes were small.
They should have been.
One early runtime used a tiny Materials Project slice, lightweight classical models and ordinary local hardware. It produced actual descriptor maps, perturbation trajectories, a candidate reduced descriptor spine and, importantly, a negative structure-aware decision rather than politely making every branch look successful.
That last bit mattered more to me than the pretty figures.
The thing could say no.
Not rhetorically.
Computationally.
A declared condition failed.
The output stayed failed.
There is a particular satisfaction in building a machine that refuses to flatter its creator.
Very healthy relationship.
Highly recommend.
And once that worked, the question got bigger.
Could it handle more data?
Another model?
Another target?
Another source?
Could the same descriptor account survive outside the first environment?
Could the experiment itself survive being interrupted?
Could I rerun it and actually tell whether the scientific result repeated, rather than simply getting another plausible-looking table?
Could the system preserve a failed replay without me or the AI quietly deciding it was probably close enough?
This is where the small runtime began turning into a programme.
Not because I had drawn a roadmap saying:
Step one: furnace.
Step two: reinvent materials assurance.
The chronology was much stupider than that.
I could not access SolidWorks.
So I built a little thermal tool.
The little thermal tool taught me I could turn my kind of specification into executable systems through GPT and Codex.
That made the coded version of the wider architecture feel possible.
Once the architecture could run, I could point it at a materials question.
Once the materials question could run, the result started disagreeing with the assumptions underneath the materials question.
And every disagreement created another piece.
Poly Bridge.
Again.
Always fucking Poly Bridge.
The strange thing is that, somewhere in that sequence, code stopped feeling like a separate skill I did not possess.
It became another material.
Language was one material.
Diagrams another.
Equations another.
Code another.
Different properties.
Different failure modes.
Use the one appropriate to the bit of the structure you are trying to carry.
That was much less intimidating.
I still did not particularly enjoy debugging.
I remain confident nobody actually enjoys debugging and some people have simply developed Stockholm syndrome.
But I no longer needed to identify as A Programmer before I was allowed to make executable things.
I needed enough understanding to specify, inspect and reject.
The rest could be collaborative.
Human plus language model plus coding agent.
Not because the AI was replacing the engineer.
Because the interface had changed where the engineer's effort went.
Less:
remember exact syntax for every operation.
More:
define what must remain true while the operation is built.
That suited me extremely well.
Possibly suspiciously well.
And it completed a loop that had started years earlier with that Year 3 report.
Show your working.
I had spent most of my life seeing the answer before I had the route another person could follow.
Then AI helped me translate the route into language.
Now that language could be translated again into software.
And software did something language alone could not.
It forced the route to exist.
Not just persuasively.
Operationally.
Inputs went in.
Something happened.
Outputs came out.
If I claimed the route did one thing and the program did another, we had a disagreement I could no longer fix with a nicer paragraph.
This was precisely what I needed next.
Because I had just spent Act V discovering that two highly articulate systems could construct an extremely convincing version of reality between them.
Now I had a way to invite something much less polite into the conversation.
Execution.
And execution, unfortunately for several of my early ideas, has terrible bedside manner.
The first proper materials programme did not give me the clean little descriptor hierarchy I had gone looking for.
Good.
That is where Geatomica actually starts becoming interesting.
Because the code worked.
Then the experiment started refusing the story.
And from there, reality got increasingly involved.
Act VII — A Slight Issue With a Zero
There is one fairly important detail that makes the end of the last act slightly less heroic.
The thermal tool fucked up.
Or, more accurately, the thermal route produced something stupid and then I looked directly at the stupid thing and accepted it.
Team effort.
I still have not actually finished the post-mortem on what happened inside the original engine. I know the number that came out, I know what I thought it meant, and I know those two things were very different. Somewhere in the route, the wall thickness ended up around 335 to 400 millimetres when the compact architecture I had in my head was more like a few centimetres. I have not yet gone all the way back through the old code and receipts to work out exactly where that happened, because shortly afterwards I got distracted by the entirely reasonable activity of accidentally building Geatomica.
I will go back to it.
This essay has already established that “I will go back to it” is not a statement to which a sensible person should attach a deadline.
But the actual mistake is extremely simple.
Three hundred and fifty millimetres is thirty-five centimetres.
I knew this.
Obviously.
Unfortunately, by that stage of the project I was very tired, very ill and apparently no longer operating a fully licensed copy of the metric system.
So I had been looking at the number and mentally reading the architecture as roughly three and a half centimetres.
Not thirty-five.
Three point five.
Lost a zero.
Nothing major.
Just approximately ninety percent of the wall.
The really beautiful part is how I found out.
Not calmly at home.
Not while doing a final check.
Not while staring at the code thinking, hold on, this looks a bit chunky.
During the presentation.
They asked me something along the lines of:
“How big do you think 350 to 400 millimetres actually is?”
And I, with the confidence of a man who has spent several weeks designing the thing, held my thumb and index finger apart.
About this much.
A few centimetres.
There was a small pause.
Then they looked at me and used both hands to show the actual distance.
About this far.
And I remember looking at it and having the entire problem rearrange itself in my head in real time.
Three hundred and fifty millimetres.
Thirty-five centimetres.
Oh.
Balls.
There are moments where you can physically feel your internal narrative receive a software update.
Externally, though, we remained calm.
Very calm.
Good thing I had spent most of my life practising that particular skill.
Inside my head there were a lot of curse words, most of them being directed with unusual precision toward myself. Outside, I basically had to go:
“Ah. Yep. That’s my mistake. I’ve been interpreting that as around 3.5 centimetres.”
Oopsie.
My fault.
Anywho.
Because what else are you going to do?
There is no credible engineering defence based on:
I simply did not realise thirty-five centimetres was quite large.
You take the L.
And, weirdly, once I had taken it, there was still a project there.
That was the part I had to pivot into during the presentation.
The actual viewport architecture was not fundamentally dependent on one sacred wall thickness. The whole point of the modelling route was that thickness could move. Material properties could move. The cartridge geometry could be changed. You could rerun the thermal problem around the state you actually wanted.
So I basically had to spend the rest of that conversation explaining:
Yes, that number is wrong.
No, the entire engineering idea has not therefore evaporated.
The cartridge still separates jobs that are better kept separate. The glazing does not need to carry every structural responsibility. The seal has a job. The clamp path has a job. The insulation has a job. The frame has a job. The thing is modular. If the furnace wall is smaller, you scale the architecture and rerun the thermal route.
The thermal model had always been useful precisely because I did not want the design welded permanently to one thickness.
So the presentation became slightly less:
Here is my beautifully resolved final answer.
And slightly more:
Ignore the absolutely enormous wall I have accidentally constructed and let me explain why the rest of this thing is actually pretty decent.
Which, to be fair, I think landed better when I could talk through it.
That has always been the annoying thing.
If you put me in a room with the actual mechanism and let me explain why something exists, answer questions, move between bits of the design and show how they connect, I am normally much better.
The fixed artefact is where life gets more entertaining.
Report.
Poster.
Slides.
Page limit.
Deadline.
Whatever amount of cognitive capacity is available while I am making them.
Whatever visual system has decided to become my enemy that week.
And suddenly something that feels very clear in my head becomes considerably less clear in the thing somebody else actually has to judge.
Which, yes, is once again basically my Year 3 report coming back for another shift.
Show your working.
I KNOW, MISS.
I AM TRYING.
There is one final insult here.
I later discovered I probably could have got the remote connection working in the first place.
The university had changed the system between my fifth and sixth year and I had simply not properly read through the changes.
Eventually I worked it out and got back onto the CAD software.
By which point I had already built the thermal tool.
So, to recap:
reading the patch notes: apparently too much admin.
building a thermal-sweeping software package from scratch because I thought remote CAD was unavailable: yeah alright then.
ADHD remains a serious condition.
The frustrating part was that I actually thought the furnace project contained some genuinely good engineering.
I still do.
It was not my best university project as a finished academic object.
I don't think the thing I finally got to show represented what I had wanted to make.
There were bits of the design that had developed faster than I managed to communicate them. The visual side never really reached where I wanted it. Some decisions made more sense in conversation than they did on the page. And then, obviously, there was the small matter of temporarily discovering a thirty-five-centimetre wall and only noticing while somebody was physically demonstrating its size to me.
Not ideal.
But from my side, the project had also become much bigger than the thing I had originally been asked to do.
I had started with a viewing-port problem.
Somewhere along the way, because I could not reliably access the normal software route, I had built a local thermal analysis tool.
That is still slightly mad to me.
I had not gone into the project thinking:
Excellent. This is the month I become a software developer.
I mostly wanted to use SolidWorks and whatever else I needed without having to drag my body into university and sit there for hours.
Instead I ended up learning that I could take a mechanism I understood, talk it through with GPT, turn the conversation into a more explicit specification, hand the implementation work to Codex, run what came back, break it, inspect it, correct it and gradually get to something useful.
That was a big deal.
Then, unfortunately, the first proper engineering system I built also gave me a live demonstration of the difference between making software and making correct software.
Quite educational.
Would perhaps have preferred the lesson in private.
The thing I like about the mistake now is that it was not a clean “AI hallucinated” story either.
That would have been easier.
Machine wrong.
Human clever.
Human catches machine.
Article written.
No.
The engine produced something silly.
Then I supplied the human verification layer.
And the human verification layer went:
yeah that seems fine.
So between us we achieved a highly governed, conversationally assisted and computationally reproducible mistake.
Innovation.
It is probably one of the most useful reminders I got during that whole period.
You can spend as long as you want worrying about whether the AI will change your instructions.
There is another failure mode where it does exactly what you allow it to do and you are the weak link.
I had spent months trying to get AI closer to my own tolerances.
Preserve this.
Do not change that.
Keep the mechanism.
Do not invent the missing history.
Carry the correction forward.
And now I had an executable system where the correction problem could come from the other direction.
The machine can be wrong.
You can be wrong.
Worse, you can be wrong together.
Very wholesome.
Human in the loop sounds great until you remember the human is also a biological organism who may be exhausted, ill, distracted, under deadline pressure and apparently unable to visually distinguish three and a half centimetres from the length of a decent ruler.
The loop inherits the loop.
That sounds obvious, but I had been so focused on making AI behave that I had not really confronted the opposite question yet.
What happens when the AI preserves my mistake?
I did not have some complete answer to that then.
Still don't.
You add checks.
Units.
Ranges.
Independent validation.
Different routes that should disagree if something has gone badly wrong.
You make important assumptions visible.
You stop letting one number silently carry an entire design state.
But ultimately, somewhere, somebody has to remain willing to ask:
Hang on.
Does this make any fucking sense?
And during the furnace project, for one important number, the answer was no.
I just got around to asking it slightly late.
There was also something frustratingly appropriate about this happening while I was struggling so much with university more generally.
That part is difficult to explain without sounding contradictory.
I could barely engage with some of the formal work I was supposed to be doing.
Attendance was difficult.
Access was difficult.
Starting things was difficult.
Finishing things in the right format was difficult.
There were periods where the ordinary administrative layer around being a student felt almost comically expensive.
And at the same time I could apparently spend a month learning enough about AI-assisted software development to build my own engineering analysis route because there was a technical problem annoying me and the normal route seemed inaccessible.
That is a very weird place to exist.
It does not feel like being “high functioning”.
It also does not feel like being incapable.
It mostly feels like having access to a bizarre set of doors where some of the tiny ones are welded shut and then, for no apparent reason, a massive industrial shutter halfway down the corridor is completely open.
You spend forty minutes trying to answer an email.
Then build software.
Normal.
I think that was part of why the dimensional mistake hurt more than it should have.
Not because one bad number destroyed everything.
It didn't.
It was because I had put so much effort into finding another route.
I had got around the access problem.
I had made the software.
I had made the model run locally.
I had built something genuinely new to me.
And then I had managed to undermine part of my own work with something so basic that explaining it almost sounds fake.
Three hundred and fifty millimetres.
Thirty-five centimetres.
Come on, man.
You do have to laugh.
The alternative is basically throwing the PC out of the window, which is an expensive response and would have created several additional problems.
So you laugh.
You say oopsie.
You explain the scalable architecture.
You keep your face straight.
You go home.
Then internally you reopen the disciplinary proceedings.
I was genuinely angry at myself.
Not in some useful reflective engineering sense either.
Initially it was more:
you absolute fucking idiot.
Wonderful.
Brilliant.
Spent all this time on thermal modelling and got taken out by Year 5 maths.
Elite.
But eventually the frustration settles down enough that the mistake becomes interesting.
Because the software itself had still changed something fundamental for me.
Before the project, I did not really think of myself as somebody who could make software.
Afterwards, whether the first tool had a silly dimensional bug or not, that belief was gone.
I had made one.
It existed.
It ran.
It produced results.
It let me scan a design space.
It let me do work I otherwise could not easily access.
And I understood enough of the workflow to do it again.
That was the actual thing I left with.
Not:
Joe has mastered thermal analysis.
Clearly not.
More:
Joe now knows how to turn one of his conceptual systems into something executable.
That is a much more dangerous piece of information.
Because I already had a lot of conceptual systems.
Chief-OS was sitting there with all its written modules.
Memory.
State.
Recovery.
Governance.
Different little pieces of the conceptual empire that had spent months existing primarily as words because words were the substrate I knew how to manipulate.
Then there were the engineering ideas.
The materials ideas.
Old machine-learning work.
Design of Experiments.
Questions I had been carrying around about what happens when you perturb a system instead of just measuring it once.
And now I had a route.
I did not need to personally translate everything directly into code.
That had been the psychological wall before.
Instead:
I understand the mechanism.
Talk it through.
Make the requirements explicit.
Use GPT to help turn those requirements into a buildable sequence.
Give the sequence to Codex.
Inspect what comes back.
Run it.
Break it.
Correct it.
Repeat.
Suddenly “this could probably be coded” was no longer an interesting sentence.
It was a threat.
Because now the obvious follow-up was:
Alright then.
Code it.
And I was already learning from the furnace that executable did not mean correct.
Which, in hindsight, is probably the best possible lesson to acquire before pointing the workflow at something more ambitious.
If the thermal tool had been flawless, I might have come away slightly too impressed by the ability to produce software quickly.
Instead, I had simultaneously learnt two things.
I can build this stuff.
And I can still be wrong.
Good.
Useful combination.
Keeps everyone humble.
The funny thing is that I still have not gone back and properly finished that little investigation.
I want to.
There is a proper answer in there somewhere about exactly how the thickness moved through the original engine and why it did not get caught earlier.
Maybe the input was wrong.
Maybe something in the model route transformed it.
Maybe I simply accepted a configuration that had drifted away from the physical object I thought I was representing.
I don't want to invent the explanation now just because one would make this paragraph cleaner.
That would be a fairly ironic thing to do at this stage of the essay.
So for now:
I know the bad state existed.
I know I accepted it.
I know the live presentation contained the world's most efficient unit-conversion intervention.
The detailed forensic report remains pending.
I got distracted.
The distraction was called Geatomica.
This is probably where the project history gets particularly ridiculous.
Because instead of finishing the furnace, fixing every visual, completing the software autopsy and then calmly moving onto the next thing like somebody with project management skills, I looked at the new capability I had acquired and basically went:
Hang on.
I can use this on the other stuff.
Of course.
I had spent months developing ideas in conversation that kept reaching the same annoying limit.
Interesting model.
Nice framework.
Could probably be tested.
Would need software.
Previously, “would need software” had been a perfectly serviceable place to stop thinking.
Now it wasn't.
I'd lost the excuse.
And the materials side was the obvious place to start because I actually knew the domain.
I had done materials science for years.
I had done machine-learning work.
I had done Design of Experiments.
I had spent my placement preparing physical evidence and learning how much work happens before interpretation.
I had just spent a final project learning, among other things, that a computational system can be internally very helpful while carrying a stupid assumption.
There was a particular question sitting around that had been irritating me for a while.
Machine-learning models tell you which features are important.
Fine.
But important in what sense?
Does the same thing remain important if you change the model?
The population?
The noise?
The split?
What happens if you remove the thing and make the model learn again?
Is the feature revealing something durable about the material problem, or something convenient about the particular route that happened to produce the result?
That was exactly the kind of itch I had never been very good at leaving alone.
And now, unfortunately, I had learnt how to scratch it.
So I started building.
Again, not with some grand plan to create a materials-assurance research programme.
At first it was much smaller.
Could I make an engine that takes a materials dataset, trains models, perturbs the setup and records how the importance structure changes?
Could I stop looking only at the best-performing model and instead look at how the relationship behaves when I mess with the conditions around it?
Could I use the machine as an experimental instrument rather than merely a prediction machine?
This seemed reasonable.
Reasonable questions are dangerous around me.
Because once the first version works, it immediately produces another question.
Then another.
Then the output needs remembering.
Then runs need comparing.
Then a crash means recovery matters.
Then two runs disagree and suddenly replay means something.
Then a result that looks good against one comparator looks terrible against another.
Then a source field exists but you realise existence is not the same thing as physical authority.
And before long the little materials experiment has acquired enough administration that Chief-OS is looking over from the corner like:
finally, somebody else understands.
This is where my relationship with university work and the self-directed work became even stranger.
The formal work still had all the things that should make work easier to evaluate.
Deadlines.
Questions.
Assessment criteria.
People whose job was literally to tell me whether I had done it well.
And I was struggling immensely with it.
The other work had none of those.
Nobody had asked me to do it.
Nobody was marking it.
Nobody was waiting for it.
There was no supervisor telling me which branch was sensible.
No research group.
No external person regularly looking at the results and saying:
yes, that is interesting.
No.
That is already known.
No.
You have misunderstood something basic.
Or even:
mate, perhaps sleep.
I was mostly alone with the work and a collection of AI systems that were extremely capable of helping me continue.
Which is an odd combination if you think about it.
Because there is no natural stopping point.
A university assignment ends because the deadline arrives.
A self-directed question ends when you are satisfied.
I have just spent seven acts demonstrating that this is not a reliable stopping criterion for me.
There was also no mark.
That is stranger than it sounds.
You spend most of education having somebody else eventually give the work an externally generated number.
65.
72.
48.
Whatever.
You might disagree with it.
You might hate the assessment.
But at least something outside your own head has happened.
With this, I could run another experiment.
Get a result.
Build another component.
Run another check.
Look at the whole thing and think:
I think this might actually be good.
And then immediately:
But how the fuck would I know?
I am the one who built it.
The AI helping me is hardly an independent reviewer.
We have already established that between the two of us we can become very convincing.
And the work had not been built to somebody else's rules.
I was not answering a question that had arrived from above.
I was following the question itself.
Something did not make sense.
Scratch.
That exposes something else.
Scratch.
That distinction collapses.
Scratch.
Before long you have a research programme where a few months earlier there had been a furnace and a very unfortunate zero.
It was liberating.
It was also deeply weird.
Because I had no idea where the ceiling was.
Maybe I had made something genuinely useful.
Maybe I had made the world's most elaborate private filing system.
Maybe half of the important-looking things were already completely standard in fields I had not yet read deeply enough.
Maybe the results were obvious.
Maybe some were wrong.
Maybe the architecture was actually solving a problem other people cared about.
Maybe I was just very good at explaining why it should.
At some point, internal confidence becomes almost useless.
Especially after you have built half the architecture around the proposition that coherent stories should not be allowed to certify themselves.
Bit awkward.
I genuinely wish I could get my Year 24 report.
Miss had the Year 3 one nailed.
Makes connections.
Likes investigating.
Needs to slow down.
Needs to think about the reader.
Needs to explain the mental method used to reach the answer.
Fair.
Can I get the follow-up please?
Something like:
Joseph continues to show enthusiasm for scientific investigation.
He has made progress in showing his working.
Unfortunately, there is now far too much of it.
Joseph should seek feedback from other adults.
That would actually be quite helpful.
Because that was becoming the next thing the work needed.
External eyes.
Not another AI telling me the architecture made sense.
Not me rereading my own documents and deciding that past-me had occasionally cooked.
People who knew the domain.
Researchers.
Engineers.
Machine-learning people.
Scientific-software people.
People who might look at one of the results and immediately see something I had completely missed.
Good.
That is information.
People who might say:
this is useful.
Also information.
People who might say:
this is bollocks.
Potentially upsetting information.
Still information.
That is where the project was heading whether I liked it or not.
I had spent years trying to make my internal structure legible enough to travel.
Now it actually had to travel.
And this time I could not be the person on both ends of the communication.
I had built enough.
The next stage required somebody else to disagree with me.
Which, given my record, was probably overdue.
Still, before I could get there, I had one immediate problem.
I had finally acquired a practical route for turning ideas into executable systems.
I had an enormous backlog of ideas.
And, for perhaps the first time, “I don't know how to code that” no longer worked as an excuse.
I had the skills now.
Time to pay the bills.
Not financially.
That part remained aspirational.
The ideas had been sitting there for months.
Now they wanted implementations.
And once I pointed the new workflow at materials data, the little question I had started with began producing answers I had not expected.
That is where Geatomica stopped being another name in the conceptual empire.
It started becoming an experiment.
Act VIII — Year 24 Report Pending
So what did all of this actually become?
Fair question.
We have travelled quite a long way from me crying in front of a maths textbook.
There have been furnaces.
Medical letters.
Hypercubes.
A conceptual government.
An AI committing administrative identity fraud.
A software career I attempted to avoid by writing approximately twenty-five documents.
A thermal-modelling package I built partly because I did not read the university patch notes.
And one live demonstration that three hundred and fifty millimetres is, in fact, quite a lot bigger than three and a half centimetres.
Strong programme.
Very interdisciplinary.
But somewhere inside all of that, the little materials experiment from the end of the last act kept growing.
Not smoothly.
Obviously.
Nothing in this essay has grown smoothly.
The original idea was still fairly straightforward.
Models use material descriptors.
Fine.
If one descriptor looks important, what happens when I change the conditions around that judgement?
Different model.
Different split.
Different population.
Noise.
Removal.
Retraining.
Does the importance survive?
Does it move?
Does something else replace it?
Are we looking at something genuinely difficult for the model to do without, or merely something that one fitted model happened to use a lot?
I liked that question because it felt like failure analysis.
Not:
Which model won?
More:
What happens if I start removing the things holding the answer up?
That is much more my speed.
So I built the early runtime around that.
And then it started giving me inconvenient answers.
This is where Geatomica became much more useful to me than it would have been if the first hypothesis had simply worked beautifully.
The first big disagreement was around feature importance.
There is a very normal thing you can do with a trained machine-learning model.
You can perturb one feature and see how much the trained model’s performance changes.
If performance falls a lot, that feature looks important.
Reasonable.
I had no particular beef with this.
The problem came when I asked what I thought was basically the same question in a more violent way.
Instead of perturbing the feature inside the already-trained model, remove the feature completely.
Then throw the old model away.
Train a new one without it.
Now ask what happened.
Those sound like close relatives.
One asks:
How much does this fitted model rely on the feature it already has?
The other asks:
How difficult is it for the modelling system to recover when that information is never available in the first place?
I expected some disagreement.
Machine learning is machine learning.
Nothing ever behaves politely enough to make a conference figure in one attempt.
I did not expect the rank agreement to come out at about 0.0394.
Basically:
hello.
We appear to be answering different questions.
The funny thing is that neither measurement was fake.
That became the important bit.
I had not discovered that ordinary feature importance was useless.
I had discovered that I had been quietly allowing one legitimate measurement to impersonate another legitimate measurement.
Fitted use.
Retrained dependence.
Related.
Not the same.
That distinction became Paper 1.
The public programme now puts the result right near the front because it captures the whole problem in one number: under the frozen Matbench route, fitted permutation importance and complete remove-and-retrain consequence had mean rank agreement close to zero.
I liked that result.
Not because it proved some giant universal theory.
It did not.
I liked it because the experiment had disagreed with the easy interpretation.
Good.
That meant the machine was beginning to earn its keep.
Then the second problem appeared.
Suppose a model beats nonsense.
Surely that is useful evidence.
Again:
yes.
And again:
not quite as much evidence as I initially wanted it to be.
For the Materials Project elasticity work, I compared real predictive models against deliberately broken versions where the target relationship had been destroyed.
Same sort of logic as asking whether somebody has genuinely learnt something or is performing at chance.
Across the matched set, 73.3 percent of the real models beat their broken comparator.
That sounds encouraging.
And it was.
Then I asked the more awkward question.
Were the predictions actually good?
Not better than nonsense.
Good.
Positive held-out predictive performance.
Same 288 cases.
Thirteen passed.
4.51 percent.
Ah.
So the models had usually learnt something.
The slightly inconvenient follow-up was that what they had learnt was usually not enough to make them useful predictors under even that fairly basic absolute criterion.
Again, both observations were true.
There was real signal.
There was weak absolute viability.
I had previously treated those as if one should naturally mature into the other.
The data declined.
That became Paper 2.
And, again, this is now exposed in the public repository rather than hidden in some flattering summary: the compendium literally headlines the transition as 73.3% -> 4.51%, then immediately says these are bounded measurements rather than universal thresholds or automatic engineering permission.
That last bit became increasingly important.
Because once you start finding distinctions like this, there is a very tempting failure mode.
Inventing an architecture after the result and then acting as though the architecture predicted the result all along.
I had become suspicious enough of my own storytelling by then to try very hard not to do that.
The questions changed because the experiments made them change.
First:
Which variables appear important?
Then:
Important under what intervention?
Then:
Does the model contain signal?
Then:
Is the signal enough?
Then:
Is the source itself good enough for the interpretation I am placing on it?
Then:
Did the result repeat?
Then:
When I say repeat, what exactly repeated?
At some point I looked around and realised the original descriptor experiment had accumulated quite a lot of paperwork.
Very familiar.
The third paper became the point where the problem moved beyond the model itself.
Because even if a result is real, useful and reproducible, you still have another question:
What does that result have permission to change?
This sounds unnecessarily legal until you put engineering on the other side of it.
A database field exists.
Fine.
Does that mean I know exactly what physical method produced it?
No.
A model runs.
Fine.
Does that mean the model is viable?
No.
A result repeats.
Fine.
Does repeating it prove the physical interpretation?
Also no.
A computational result can be completely genuine and still be several evidential steps away from:
build the thing.
That became the third paper.
And this is where all the weird life stuff and the weird AI stuff started turning up inside the science whether I intended it to or not.
Not because fibromyalgia contains an epistemology of materials informatics.
Please do not put that in the abstract.
But because I had already spent years becoming extremely sensitive to the difference between:
something happened
and
what am I allowed to infer from the fact that it happened?
I was still working while my body deteriorated.
That did not mean the working pattern was sustainable.
I could still produce an assignment.
That did not mean I had reliable access to the underlying ability.
The AI could reconstruct a perfectly coherent project history.
That did not mean the history had happened.
Now a model could beat nonsense.
That did not mean the model was good.
Same offence.
Different workplace.
By then I had also become increasingly obsessed with replay.
Not merely:
Can I run it again?
But:
What exactly am I claiming repeated?
Same scientific plan?
Same cohort?
Same splits?
Same numerical values?
Same ordering?
Same files?
Same bytes?
Those are different levels of sameness.
This becomes very boring very quickly.
Excellent.
Boring is where I was starting to feel safest.
The Spatial work eventually produced one of the sillier-looking numbers in the programme.
A fresh run and replay comparison across 927,361 numeric cells.
Maximum numerical difference:
zero.
Every cell exactly matched under that comparison.
Lovely.
Except the scientific conclusion itself contained both a positive result and a conditional one.
The replay did not improve the science.
It preserved it.
That distinction matters a lot to me.
If the original result says:
bulk behaviour supports this claim under these conditions
while shear remains conditional,
then a good replay should not return:
Congratulations, everything now passes because we tried twice.
It should reproduce:
bulk passed.
shear remains conditional.
Which is what happened.
The repository now displays the 927,361 exact comparisons next to the other two disagreements, not as “therefore the whole programme is true”, but as an example of exact replay coexisting with bounded authority.
This is the point where the thing started looking enough like a research programme that I became slightly uncomfortable calling it my little project.
There were now three separate papers.
A larger monograph pulling the development together.
A technical companion describing the runtime and industrial interpretation in more detail.
Protocols.
Evidence tables.
Replay material.
Claim boundaries.
Reproduction routes.
And eventually the bit I had spent this entire essay accidentally training myself to care about:
something another person could actually inspect.
So I started building the public release.
This created another problem.
I wanted people to be able to look under the bonnet.
I did not want to publish the entire production engine I had spent months building, because I am also trying, at some point, to construct a consultancy and software business rather than becoming the world's least solvent open-source foundation.
There is a balance there.
Show enough that the scientific claims do not depend on:
trust me bro.
Do not give away every reusable bit of the machinery that might eventually pay for food.
Particularly because food has already featured in this essay as an unreliable luxury.
So the public release had to obey the same rule the research was trying to enforce everywhere else: show enough that somebody can check the claim, without pretending the public package contains things it does not.
That mattered for two reasons. Scientifically, I did not want the whole thing resting on trust me bro. Commercially, I also did not want to put the reusable Systemica and Geatomica machinery on GitHub and become the world's least solvent open-source foundation. Food had already been inconsistent enough without turning the potential software business into a charitable donation.
So the main repository became a bounded research compendium. Papers. Monograph. Technical companion. Evidence. Protocols. Reproduction. Release audit. Claim boundaries. Enough of the route for somebody else to inspect what I am actually claiming, while the reusable orchestration, memory, recovery and deployment machinery stays outside it.
The important bit was not the folder structure. It was what happened when the public package reached something it did not contain.
Earlier in the project, a missing object was exactly where an AI could become dangerously helpful. It could see the shape of what belonged there and reconstruct something convincing enough to look historical.
I now wanted the opposite behaviour.
The smaller reproduction scripts therefore do one modest job: reconstruct the reported readouts from the accepted public tables. They do not fit the models and they do not quietly become the private runtime because that would make the release sound more reproducible than the object actually was.
Reproduce the readout.
Re-run the scientific fitting.
Related.
Not the same thing.
For the second job I made a separate scientific verification rerunner. That one actually recomputes representative claim-bearing kernels from Matbench, Materials Project and the Spatial GIL route.
And because I had already learnt what happens when a system knows the answer it is supposed to recover, the expected result is deliberately kept out of the computation. compute does the fitting and writes a receipt first. verify only loads the frozen reference afterwards and asks whether the new result agrees.
The answer does not get to sit in the room while the exam is happening.
Then there is full.
If you ask the public rerunner for a complete publication campaign that has not actually been bound into the package, it does not look at the surrounding structure and go, boss, I basically know what you mean.
It writes:
FULL_PUBLICATION_CAMPAIGN_NOT_BOUND
and refuses.
That might be one of my favourite outputs in the whole programme.
Act V was a machine seeing a missing scientific object and helpfully completing the pattern.
This one sees the missing object and says:
No.
You haven't given me that.
Go away.
Growth.
The bounded real-data reruns did reproduce their frozen reference results exactly to 1e-12 on the reference workstation, with 1e-10 declared as the portable scientific acceptance tolerance. Useful evidence. Still not external replication. Still not peer review. Still not physical qualification. The same result happening again means the same result happened again. It does not mean reality has signed the rest of the paperwork.
The main compendium carries the same attitude in a tiny file called CLAIM_BOUNDARY.md. Its job is basically to stand beside the exciting results and tell them where they are no longer allowed to travel.
No universal descriptor hierarchy.
No automatic engineering deployment.
No pretending local replay is external validation.
Imagine telling the version of me from a few years ago that I would eventually become emotionally attached to a document whose main purpose was reducing the maximum permitted excitement.
Character development.
And this is roughly where I am now.
Which is a strange sentence.
Because I still do not completely know what “this” is.
I know what I built.
Mostly.
I know what the experiments did.
I know what the papers claim.
I know which bits are held back.
I know the software runs.
I know the public rerunner is separate from the production estate.
I know the evidence package contains enough for somebody else to start checking the work rather than relying entirely on the essay you are currently reading.
What I do not know is the thing education normally tells you eventually.
Is it good?
That is a surprisingly difficult question when nobody assigned the work.
University gives you a bizarre amount of external structure even when you are terrible at using it.
Here is the module.
Here is the learning outcome.
Here is the deadline.
Here is the page limit.
Here is the person marking it.
Eventually somebody gives the thing a number.
Maybe you agree.
Maybe you don't.
Still:
external event.
Geatomica did not have that.
Nobody told me to build it.
There was no supervisor meeting every Tuesday.
No research group.
No PhD cohort.
Nobody saying:
this branch is useful.
This branch is nonsense.
This already exists.
This needs another control.
Stop naming things.
Go to bed.
I had AI.
Which, as we have established, is not exactly an independent reviewer of whether the things we have spent six hours enthusiastically constructing together are secretly brilliant.
And I had myself.
Even worse.
So there is something deeply strange about sitting with a research estate like this and thinking:
I think this is actually pretty good.
Then immediately having to ask:
Compared to what?
According to whom?
Because if I have learnt anything from my own work, it would be quite embarrassing to finish by giving myself scientific authority on the grounds that I found my own story persuasive.
No.
Not having that.
The work has to leave the room.
That is partly why I am making the public package as inspectable as I reasonably can without throwing the whole private engine onto GitHub.
Also I cannot afford to throw much of anything.
Compute is expensive.
Hardware is expensive.
And if I literally tried to throw the PC, my back would probably go out.
Not out.
Out out.
Respectfully, I do not swing that way and I see no reason my spine should be making independent lifestyle decisions.
So the PC stays on the desk.
Very advanced IP strategy.
And now the work goes outside.
For anyone in materials informatics, scientific machine learning, engineering assurance, reproducibility, scientific software, or simply anyone who has made it this far and would like to check whether I have been chatting shit:
please do.
The main public research compendium contains the three papers:
Importance Is Not Necessity.
Better Than Null Is Not Yet Useful.
Available Evidence Is Not Engineering Authority.
It also contains the Geatomica flagship monograph, technical companion, protocols, selected evidence, reproduction routes and explicit claim boundaries.
And if you want something smaller that actually recomputes representative scientific kernels rather than merely reconstructing reported summaries, there is a separate Scientific Verification Rerunner.
That package is intentionally standalone and publication-specific. It does not need the private Systemica or Geatomica runtime to perform its bounded checks.
So:
here is the work.
At some point it stops being mine to grade.
Audit it.
Break it.
Tell me where I have misunderstood something.
Tell me which distinction is already completely standard under a name I have somehow failed to find.
Tell me where the statistics need strengthening.
Tell me where a claim is too strong.
Tell me where I am being too cautious.
Tell me if something is genuinely useful.
Good feedback is information.
Bad feedback is information.
Some information is admittedly more pleasant to receive.
But still.
Information.
I have tried very hard to show the working.
Which feels like the correct place to return to the beginning.
Year 3.
Eight years old.
Apparently good at making connections.
Enjoyed scientific investigation.
Needed to slow down.
Think about the reader.
Explain the mental method used to reach the answer.
I would really appreciate my Year 24 report right about now, miss.
Just something brief.
Joseph continues to make strong connections across subjects.
Joseph has demonstrated enthusiasm for independent investigation.
Joseph has made progress in explaining his working.
Unfortunately, Joseph has now generated an unreasonable quantity of working.
Joseph should seek feedback from other adults.
Fair.
Because that is the weirdest part of the whole thing.
I am twenty-four.
I built most of this alone.
Not “alone” in the stupid sense where I pretend AI did not contribute enormous amounts of labour.
It did.
That is half the article.
But there was no established research team directing the questions.
No laboratory handing me the programme.
No supervisor defining the scope.
Mostly it was me, my phone, my PC, some tunes, what I will diplomatically describe as a herbal-remedy vapour, and vibes.
That is, genuinely, how a disturbing proportion of this estate was made.
Sometimes the PC was doing the compute while I was doing the actual thinking through my phone because holding a phone wherever my body had decided to exist that day was substantially easier than sitting neatly at a desk pretending to be a normal research environment.
Very sophisticated laboratory.
Strong facilities budget.
And all of this while simultaneously struggling enormously with the degree that was supposed to be evidence I knew how to do materials engineering.
That still messes with my head slightly.
There were days where I could not make myself engage properly with university work.
Then I would end up working on something like scientific replay architecture until my body forced the issue.
There were formal assessments where my output looked worse than things I understood less deeply.
Then there was this unmarked thing in the background getting larger because nobody had told it to stop.
It is a very odd relationship with ability.
I still do not think the useful conclusion is:
secret genius misunderstood by education.
Absolutely not.
I have provided multiple counterexamples to that thesis within the last few thousand words.
I forgot how centimetres work.
The point is stranger.
Ability is not one number.
Access matters.
Context matters.
Interest matters.
State matters.
Translation matters.
Cost matters.
Whether somebody else can inspect the route matters.
And sometimes the place where you perform worst is the place explicitly designed to measure how good you are.
Annoying.
But there we are.
Looking back, the architecture did not really emerge because I wanted to build a grand system.
It emerged because I kept meeting the same class of problem in different clothes.
The visible result was not enough.
The route mattered.
Then the route was not enough.
The state mattered.
Then state was not enough.
History mattered.
Then reconstructed history turned out not to be history.
Then evidence existed but did not automatically support the stronger claim.
Then a model contained signal but was still useless.
Then a result repeated exactly but remained conditional.
Then software did exactly what I let it do while carrying a ridiculous wall thickness.
Every time two neighbouring things looked similar enough to collapse into one another, something eventually slapped my hand.
Fair.
This is why I still like the word HOLD.
For most of my life my main operational doctrine was:
KEEP CALM AND CARRY ON.
Very British.
Very useful.
Got me surprisingly far.
Also almost certainly responsible for several maintenance issues.
The systems I ended up building added another possibility.
Sometimes you carry on.
Sometimes you recover.
Sometimes you change the route.
And sometimes you have enough information to say:
not yet.
HOLD.
Do not fill the gap.
Do not upgrade the claim.
Do not spend the money.
Do not fabricate the component.
Do not tell yourself that because you managed yesterday you therefore possess infinite tomorrow.
That is probably the most useful thing the whole conceptual empire ever taught its sole civil servant.
Naturally, I am still not particularly good at following it.
As I write this, I am sitting between two resit exams I need to finish my Master's.
I have done the first one.
I have not started revising for the second.
Instead, I have spent the afternoon writing an essay about how a lifetime of getting sidetracked, building elaborate alternative routes and struggling to do the thing I am technically supposed to be doing somehow turned into a materials research programme.
I got sidetracked.
Again.
Hard to believe, really.
Miss, I think the Year 24 report writes itself.
Final note for researchers, engineers and anyone who wants to check my homework
If you work in materials informatics, machine learning, engineering assurance, scientific software, reproducibility, evidence governance, or you are simply the kind of person who sees a large technical claim and immediately wants to poke it with a stick: please do.
The essay is the story of how I got here.
The technical work is where the claims have to stand on their own.
The public programme is currently separated into three scientific papers, a technical companion, a larger programme monograph, a public research compendium, and a standalone scientific verification rerunner.
Paper 1, Importance Is Not Necessity, asks whether fitted feature importance should be treated as evidence of descriptor necessity. It compares fitted permutation importance against complete remove-and-retrain consequence.
Paper 2, Better Than Null Is Not Yet Useful, asks whether beating deliberately broken controls is enough to establish predictive viability, using the Materials Project elasticity experiments.
Paper 3, Available Evidence Is Not Engineering Authority, deals with what happens after evidence exists: availability, admission, replay, scientific memory, bounded authority and what the evidence is actually allowed to authorise.
The technical companion documents more of the runtime and industrial interpretation: evidence states, execution behaviour, recovery, replay and the machinery underneath the scientific claims.
The Geatomica flagship monograph is the wider programme account. It connects the experiments, the development path and the bounded engineering applications without pretending every early hypothesis survived intact.
The main GitHub compendium carries the papers, monograph, technical companion, evidence, protocols, claim boundaries, release audit and minimal reproduction routes. The separate scientific verification rerunner recomputes representative claim-bearing kernels without requiring the private production Systemica/Geatomica runtime. Its compute and verify paths are separated, and where a full publication campaign is not bound into the public package, the software refuses rather than reconstructing one.
So, genuinely:
here is the work.
It is now partly yours to audit.
I have tried to expose the route, preserve the ugly results and make the distinction between what I observed and what I think it means as clear as I can.
As you all know by now, showing my working has historically not been my strongest area, so feedback is very welcome.
Good feedback.
Bad feedback.
“This is useful.”
“This already exists.”
“This assumption is wrong.”
“This result needs another control.”
“You have explained this terribly.”
All information.
Some signals are just considerably more enjoyable to receive than others. ;)
Main public research compendium: https://github.com/JosephMaxwell02/materials-ml-failure-analysis
Standalone scientific verification rerunner: https://github.com/JosephMaxwell02/materials-ml-scientific-verification-rerunner
Archived research compendium (Zenodo): https://doi.org/10.5281/zenodo.21919563
Archived scientific rerunner (Zenodo): https://doi.org/10.5281/zenodo.21939447