Capability deep-dive · deliberation rooms
You added a second opinion to the model. Can you show it read anything the first one did not?
Most “multi-agent” configurations hand every agent the same file, run a few rounds of argument, and take the majority. The published measurement on that configuration is not flattering, and an engineer evaluating a committee of models should start from it rather than from a diagram.
Start with the result that argues against rooms
Two independent author groups report the same finding: vanilla multi-agent debate often underperforms simple majority voting at higher compute cost, and can degrade below a single agent. The two mechanisms they name as missing are diversity of initial viewpoints and explicit, calibrated confidence. Everything below is those two mechanisms turned into refusals, because a documented convention decays and only a mechanism holds.
Throughput per chair is a projection with no pilot behind it. We would rather print the word than a plausible number.
The seat contract, in four clauses
A seat argues one evidence slice
The evidence admitted for a decision is partitioned into a kernel every seat sees — the notification, the schedule, the loss description — and disjoint slices, one per seat. Overlap above the configured bound does not raise a warning. It refuses the convening, with a machine code, because a room where every seat read the same file is the configuration the measurement above convicts, wearing a room’s costume.
A seat posts a probability, not an opinion
Each seat emits a probability vector over the room’s outcome alphabet, scored with a proper scoring rule and decomposed into reliability, resolution and uncertainty. A seat that is confident and wrong is measurably distinguishable from a seat that is uncertain and right, and only one of those is a benching signal.
One seat exists to puncture the consensus
The Red-Queen seat does not argue an outcome. It attacks the derivation — the provenance of an artefact, an unstated assumption, a substituted default — and each attack carries a severity and a puncture probability that the survival check consumes before anything reaches the chair.
No seat rules
A seat can conclude. It cannot release. The ruling belongs to a named human chair at a durable waitpoint that survives process, pod and node death, is idempotent with respect to effect, and writes the ruling event before any effect is released. No record, no effect.
rooms, typed messages, persona invocation(shipped)live room relay(shipped)approval waitpoints(shipped)never-dropped audit emitter(shipped)
A bible is a version, and promoting one is itself a decision
A persona bible is not a prompt in a text box. It is a versioned artefact with declared competence over artefact types, anchored to the clauses it may reason about, carried in the database and in the memory graph together, and evolving through recorded revisions rather than by edit-in-place.
A bible earns a live seat by beating the incumbent on a sequestered drill corpus — a calibration gym, with leakage rules, a champion–challenger record and an auto-benching rule for a seat whose measured skill decays. Promotion is a governed change that carries its own certificate, so “which version of which persona argued this case” is a question the record answers rather than a question the vendor answers.
evolving bible pattern(shipped)persona definitions with system prompt and memory config(shipped)the gym’s insurance drill corpus(designed, not built)
What reaches the chair, and what the dissent map is not allowed to hide
Dissent is computed as a divergence matrix over the seats’ final-round verdicts, weighted by measured calibration — because disagreement from a seat with no measured skill is noise, and a metric that treats it as signal routes noise to the chair. The unweighted spread is reported beside it, so “the disagreement is coming entirely from the weakest seat” is visible rather than hidden inside one index.
THE MAP DEGRADES BEFORE IT LIES
The dissent map is a two-dimensional embedding of a higher-dimensional disagreement, so it carries its stress statistic on its face. Above the stress limit the map is replaced by the raw divergence matrix rather than drawn. A dissent map that always looks clean is a dissent map that is lying about its own geometry.
Where a calibrated minority’s pooled evidence outweighs the majority’s, the room does not flip the verdict. It raises a flagged minority with the terms of the inequality shown, and puts the objection in front of the accountable human — which is the whole point of having one.
Where the room fails closed
REFUSED
DEADLINE_EXCEEDED
The chair’s waitpoint reached its deadline without a ruling. The room produced no verdict, and no effect was released.
A deadlocked room with a chair timeout has an honest terminal state, and it is not “proceed”. Deadline expiry escalates up the authority matrix as a new waitpoint; it never releases the decision on its own. The same law makes the reverse true: the ruling event is written durably before the effect happens, so a ruling that could not be recorded is a ruling that did not take effect.
The honest limit
WHAT THIS DOES NOT DO YET
- The chamber is the contested lane, not every decision. Throughput per chair is unmeasured [P: to be calibrated]; straight-through decisions skip the room, but they never skip the gate or the record.
- Tiers 3–5 of the dispatch ladder. Tiers 1 and 2 run in production.
- The calibration gym’s insurance drill corpus, and every threshold in the slice-admission predicate — the overlap bound, the minimum slice size and the competence floor are all [P: to be calibrated].
- A room owes a baseline comparison permanently. Until it beats a calibration-weighted pool on your own decided cases, what you are buying is calibration, and we will say so.
The throughput objection is the right first objection, and it has a design answer rather than a reassurance: the room is where contested decisions go. The routing decision that sends a case to a room instead of down the straight-through lane is itself a logged event, so the proportion of your book that ever sees a committee is a measurable property of your own configuration and not a claim of ours.
What this connects to
Convening triage
What decides that this case needs a room, which chair, and which autonomy tier — and why that routing decision is logged.
Cross-model gates
What the room’s conclusion has to survive before anything is released, and what the agreement predicate does not catch.
Evidence fabric
Where the thread goes: one governed event stream, compiled per regulator rather than reconstructed per audit.