Capability deep-dive · portfolio stress
The board asks what the tail costs. You hand over one number with no way to reproduce it.
The spreadsheet that produced it has been edited since, the assumptions live in a different file, and the person who ran it has moved teams. Nobody doubts the number because nobody can check it — which is a different thing from the number being right.
Reproducibility is the claim. The result is not
The engine is seeded and pure. The twin’s scenario dispatch does not yet record a seed or a run identity, which is exactly the gap between a run and an admissible run.
SAME SEED, SAME RUN, OR THE DOSSIER REFUSES
A run whose seed does not reproduce is inadmissible, and the dossier compiler refuses to cite it. That is the inversion this page rests on: the headline claim is about reproducibility, not about results. A distribution looks like evidence, which makes a chart the most dangerous object on a marketing site — so where a shape is drawn here at all it is drawn as a shape, every displayed value marked as uncalibrated inside its own frame.
An immutable snapshot plus a lattice of reversible deltas
Reversibility is in the object graph, not in a flag
A scenario is a row that references an immutable base snapshot; its results land in their own table; deletion is soft and preserves results. The base is never mutated, so a stress run cannot damage the thing it was run against. Nobody has to remember to set anything.
A scenario on an unfinished snapshot is refused
A snapshot that is still being built has an empty state, and a scenario over it would produce, in the source’s own words, garbage results that look like a legitimate outcome. That case throws today rather than computing. It is the shipped refusal this whole architecture is an extension of.
The engine has no domain in it
The engine’s entire coupling to its subject is one function type: samples in, a number out. Everything else is domain-free. Swapping a revenue driver for a loss driver is a closure swap, not a rewrite — which is why the maths is reusable and why nothing about the maths makes a catastrophe model appear.
A scalar delta is not a capital answer
The quantity a stress cycle reports is a distribution, and the delta between two distributions is not the difference of their means. The honest minimum is a percentile band with its tail support — a band rendered without the number of tail samples behind it is false precision, and the comparison surface is built so that it cannot be.
seeded correlated Monte-Carlo(shipped)snapshots, scenarios, comparisons, back-test store(shipped)run identity on the scenario dispatch(designed, not built)the governed stress cycle and its board projection(designed, not built)
What a governed stress cycle adds, and what is missing from the code today
The cycle is a fold over an append-only thread: freeze the base snapshot, register the models and licences in scope, seal the stress specification before any run, dispatch the runs with a recorded identity each, compile results into bands with tail support, gate the reported figures across models, deliberate with the adversarial seat, and rule — with a named chair, reasons that cannot be a template default, and a record of what was seen and which alternatives were shown. The board pack is then a projection of that thread rather than a parallel document, and no figure may appear in it that no event produced.
FOUR ACCOUNTABILITY HOLES IN OUR OWN CODE, NAMED
- The person who resolved an anomaly is not recorded — the parameter is present in the signature and unused.
- The person who executed an intervention is not recorded at all; approval is attributed, action is not.
- Approval is a status field, not a gate: there is no waitpoint, no reasons capture and no certificate behind it.
- A snapshot comparison has no author, so a delta that will sit in a board pack has no attributable requester.
None of these is a defect in a revenue twin. All four are disqualifying in a capital object, and all four are cheap to close because the object graph is already right. A compendium about accountability that read its own code selectively would forfeit the argument, so they are printed here.
What a back-test can and cannot validate
Insurance accuracy is not a time series, because the label arrives late and arrives continuously. It is a matrix of origin cohort against maturity lag, and a per-period accuracy row is a diagonal of that matrix — which is precisely the diagonal that mixes immature and mature cohorts into one number. What belongs in the cells is a distributional test on the probability integral transform, and a score in loss units for anything a committee will read, rather than a percentage error on a capital forecast.
And the bound: nothing in a back-test validates a one-in-two-hundred. The table validates the body of the distribution, the frequency, the vulnerability by cohort and the aggregate — and it has to say so on its own face. A calibration entry also has to become an entry rather than an edit, because a back-test needs to know which parameter value was live for a cohort at the time, and today calibration overwrites.
Where a stress run fails closed
REFUSED
CHAIN_DIVERGENCE
A replay of this run under its recorded seed produced a different result. The run is inadmissible and the dossier compiler refused to cite it.
A capital figure that cannot be reproduced is an assertion, and an assertion is exactly what a validation regime is trying to eliminate. The submitted seed is not the identity — it is coerced before use, so distinct submissions can collapse to one stream — so the record stores the coerced seed with the code and toolchain digests beside it. Two neighbouring refusals share the shape: a run with no deadline result is a loud machine-coded refusal rather than an indefinite wait or a silently substituted prior run, and a comparison across a changed driver set refuses rather than differencing two incomparable objects.
The honest limit
WHAT THIS DOES NOT DO YET
- The catastrophe financial chain — the exposure object, vulnerability and damage curves, deductible and limit application, exceedance-probability curves and any stochastic event catalogue. What exists is peril physics and the analytics shell, not a catastrophe model.
- The UNO→simulator dispatch seam. The simulator is deployed and the orchestrator is deployed; the executor that lets one drive the other does not yet do so.
- Exposure sums are not modelled losses. A total insured value inside a footprint is an exposure figure; running it through a correlated engine does not turn it into a modelled loss, and no arrangement of this page implies that it does.
- The governed stress and capital cycle itself is designed, not built. Capital statements would run as reversible overlay runs with a chair’s sign-off; today the twin has the object graph and not the gate.
- Back-testing coverage against a real book is unmeasured [P: to be calibrated], and the accuracy store is a period series where insurance needs a cohort-by-maturity matrix.
- The twin scopes on the tenant axis alone. Threading the owner and business-unit axes through its repositories is named work — the same predicate every other repository already inherits.
What this connects to
Validation evidence
Where a reproducible run becomes a dossier — and the refusal that stops an irreproducible one being cited.
Peril simulation
The deterministic side of the same problem, and the quarantine that makes a scenario reversible.
Deliberation rooms
Who signs a capital statement, what a substantive sign-off records, and the measurement that would catch a ceremonial one.