Skip to content
HONESTASDecision-Evidence Operating System

Capability deep-dive · portfolio stress

The board asks what the tail costs. You hand over one number with no way to reproduce it.

The spreadsheet that produced it has been edited since, the assumptions live in a different file, and the person who ran it has moved teams. Nobody doubts the number because nobody can check it — which is a different thing from the number being right.

Reproducibility is the claim. The result is not

≥ 10,000iteration floor in the correlated engine, which is seeded and pure — no clock, no input or output, no logging[O]
copulahow the engine correlates its draws before any domain model is called; the correlation specification is the caller’s[O]
no seed fieldon the twin’s own scenario dispatch today — run identity is a named build item — [P: to be calibrated][P]
9.52%the power of a ten-year back-test at the one-in-two-hundred level against a model wrong by a factor of two[D]

The engine is seeded and pure. The twin’s scenario dispatch does not yet record a seed or a run identity, which is exactly the gap between a run and an admissible run.

An immutable snapshot plus a lattice of reversible deltas

Reversibility is in the object graph, not in a flag

A scenario is a row that references an immutable base snapshot; its results land in their own table; deletion is soft and preserves results. The base is never mutated, so a stress run cannot damage the thing it was run against. Nobody has to remember to set anything.

A scenario on an unfinished snapshot is refused

A snapshot that is still being built has an empty state, and a scenario over it would produce, in the source’s own words, garbage results that look like a legitimate outcome. That case throws today rather than computing. It is the shipped refusal this whole architecture is an extension of.

The engine has no domain in it

The engine’s entire coupling to its subject is one function type: samples in, a number out. Everything else is domain-free. Swapping a revenue driver for a loss driver is a closure swap, not a rewrite — which is why the maths is reusable and why nothing about the maths makes a catastrophe model appear.

A scalar delta is not a capital answer

The quantity a stress cycle reports is a distribution, and the delta between two distributions is not the difference of their means. The honest minimum is a percentile band with its tail support — a band rendered without the number of tail samples behind it is false precision, and the comparison surface is built so that it cannot be.

seeded correlated Monte-Carlo(shipped)snapshots, scenarios, comparisons, back-test store(shipped)run identity on the scenario dispatch(designed, not built)the governed stress cycle and its board projection(designed, not built)

What a governed stress cycle adds, and what is missing from the code today

The cycle is a fold over an append-only thread: freeze the base snapshot, register the models and licences in scope, seal the stress specification before any run, dispatch the runs with a recorded identity each, compile results into bands with tail support, gate the reported figures across models, deliberate with the adversarial seat, and rule — with a named chair, reasons that cannot be a template default, and a record of what was seen and which alternatives were shown. The board pack is then a projection of that thread rather than a parallel document, and no figure may appear in it that no event produced.

What a back-test can and cannot validate

Insurance accuracy is not a time series, because the label arrives late and arrives continuously. It is a matrix of origin cohort against maturity lag, and a per-period accuracy row is a diagonal of that matrix — which is precisely the diagonal that mixes immature and mature cohorts into one number. What belongs in the cells is a distributional test on the probability integral transform, and a score in loss units for anything a committee will read, rather than a percentage error on a capital forecast.

And the bound: nothing in a back-test validates a one-in-two-hundred. The table validates the body of the distribution, the frequency, the vulnerability by cohort and the aggregate — and it has to say so on its own face. A calibration entry also has to become an entry rather than an edit, because a back-test needs to know which parameter value was live for a cohort at the time, and today calibration overwrites.

Where a stress run fails closed

The honest limit

What this connects to

Validation evidence

Where a reproducible run becomes a dossier — and the refusal that stops an irreproducible one being cited.

Peril simulation

The deterministic side of the same problem, and the quarantine that makes a scenario reversible.

Deliberation rooms

Who signs a capital statement, what a substantive sign-off records, and the measurement that would catch a ceremonial one.