Evidence & T echnical Specimens The working artefacts behind Systemica Engineering MAY 25 The development of stronger proofing for the engine is ongoing. The current work includes front-end cleanup, larger sample testing from the Materials Project, runtime proofing, practical papers, images, minimal viable demos and video walkthroughs. These are not all included in this first public specification. They will be finalised and publicised as the runtime becomes cleaner. This page is a first explainer. Its purpose is to show what Systemica Engineering is building, what technical artefacts already exist, and how the current evidence supports the direction of the work. The core idea is simple: technical claims should earn their confidence before people act on them. Systemica Engineering is being built around evidence packs: structured outputs that show assumptions, constraints, model behaviour, rejected options, residual risk and the level of confidence a claim has earned. This is not a claim of certification. It is not a replacement for domain expertise, high-fidelity simulation, physical testing or peer review. It is a pre-validation and audit layer. 1. GIL/EDK: Behavioural Feature Auditing GIL/EDK is the current model-audit runtime. It was developed around a simple problem in machine learning: feature importance is often treated as explanation too early. A feature or descriptor may rank highly in one fitted model, but behave unreliably when the data slice, model family, split, perturbation level or audit condition changes. In scientific or engineering settings, that matters because feature interpretation can influence physical follow-up, design decisions or research direction. The GIL/EDK papers reframe feature importance as a behavioural object: static rank → perturbation trace → behavioural audit The descriptor-stability paper states the aim directly: the audit records descriptor drift, descriptor variability, recurrence, negative-control behaviour, tier-removal effects and phase-density occupancy, not to prove causality or physical truth, but to ask which descriptors remain reliable enough under stress to deserve follow-up. The broader feature-auditing paper makes the same point in general machine-learning terms: a static ranking is a snapshot, while an audit trace is a behavioural object. Current outputs include: descriptor drift and variability tables baseline versus audit ranking comparisons dynamic perturbation sweeps dropout threshold plots phase-density projections barcode checkpoint tables claim-ceiling style action labels The important output is not just a plot. It is a decision-support layer. The barcode states translate model behaviour into action guidance. For example, descriptors may be labelled as invariant_core, stable_influence, conditional_stable, fragile_proxy, volatile_collapse or interaction_dominant, with each state mapped to a different follow-up action. This makes the audit practical: stable descriptors can be retained for cautious follow-up conditional descriptors need recurrence or sweep testing target-adjacent descriptors can be useful predictors but weak explanations fragile proxies should not be over-read volatile descriptors should not be used as explanatory evidence The current materials proof uses a shear-modulus dataset as the origin artefact. It is explicitly treated as proof-of-concept and validation-track evidence, not broad materials-property validation. The paper notes that the current evidence layer d