Pharos

Constraint-first AI governance

The system that builds a decision is never the system that certifies it.

A working method for putting high-stakes AI decisions through plural human-modelled review — and proving, to an auditor who trusts no one, that the review was real and not theatre.

Construct

Roles built from whole careers

Each council seat is distilled from a real expert's complete body of work into deterministic decision rules. The deeper the source, the deeper the seat.

built by the practitioner

Falsify

Roles tested by an outsider

Before any seat can vote, an independent protocol pushes it onto cases it was never trained on — to see if it reasons like the expert, or merely repeats them.

tested by an external instrument

Most "AI governance" fails because the same hand that builds the reviewer also grades it. This method keeps building and certifying on opposite sides of a hard line — and that single separation is what an auditor can actually verify.

1 How one decision is produced

From first triage to a signed, reviewed decision.

Eight stages. Tap any stage to open it. The two red stages are the certification gate — the part that makes the rest trustworthy.

2 The cabinet map

A full government in seats — but only the relevant rooms ever convene.

Each domain of statecraft is a cabinet of experts, modelled from real careers and certified before it can vote. The council can field thirty or more lanes. It almost never does. Like an attention layer in a language model, only the seats a case actually touches are activated — the rest stay dark, costing nothing.

sparse activation

A routine procurement note lights up four cabinets. A cross-border data ruling lights up nine. A sovereign-scale AI decision convenes the full chamber. The method scales its cost to the consequence of the question, not the size of the roster — so plurality is always available but never wasteful.

3 Three decisions, three depths

The same method, convening more of itself as the stakes rise.

Triage sets the depth. Tap a case to see which cabinets convene, which stay dark, and where the certification gate sits. Watch the chamber fill as the consequence grows.

4 Why it holds up

The difference between a council that looks plural and one that is.

A real-world governance product was reviewed against this method. It had built the reviewers. It had never built the part that proves the reviewers are real.

Plural by appearance

what most systems ship
  • The builder of the council also certifies that its dissent is genuine.
  • A dissenting seat is present in the roster but holds no power to change the outcome.
  • Status is claimed in prose: "deployed," "production," "signed."
  • Continuity is preserved by naming it, not by testing it.

Plural by structure

what this method requires
  • + An external protocol certifies every seat — the builder never grades their own work.
  • + A seat that fails certification is demoted, not seated. Dissent earns its vote.
  • + Status is bound to receipts: every claim carries a verifiable trail.
  • + Continuity is re-tested on the finished decision before it can publish.
5 The certification gate, in detail

Two ways a modelled expert can be fake — and the test for each.

Determinism makes different AI models converge on similar answers, so swapping models proves nothing. Only pushing a seat off its source material reveals whether it has real judgement underneath.

Failure mode 1
Miscasting

The seat carries the right knowledge but voices it through the wrong cultural lens — an Eastern or European perspective rendered with a Silicon Valley accent. The fix is matching the model's own grounding to the seat's domain, not chasing the "best" model.

Failure mode 2
Nominal continuity

The seat recites the expert's known positions but cannot reason as they would on a case the source material never covered. It cites continuity instead of perceiving it. This is the failure that looks fine until the decision actually matters.

6 Inside the contrarian cabinet

Dissent built as three instruments, each held to evidence.

The seat that fights groupthink is not a single sceptic. It is three bounded tools, and each one is forbidden from inventing a cleverer story than the record supports. The discipline is the point — an unbounded contrarian is just another way to lose the argument.

attacks the framing

Reframing lens

Asks whether the decision is solving the right problem at all. Catches false binaries, stated-preference worship, and willpower-heavy fixes where a change to the system would work better. It reframes; it never persuades.

the behavioural-reframing instinct, made operational

attacks the wording

Adversarial communications → in depth

Runs only as a controlled pair — one lane models the most hostile reading a decision will face, the other hardens the wording without hiding any tradeoff. The hostile lane is never allowed to ship alone.

simulate the attack, then defend honestly

stresses, never votes

Narrative stress

Pressures the story a decision tells about itself, then abstains. It holds no weight in the tally by design — its licence to say the unsellable thing comes precisely from its inability to change the outcome.

the court fool, given a seat but not a vote

integrity gate

The contrarian stops itself

Each instrument carries a hard stop. If the reframe depends on invented evidence, if the rewrite hides a real blocker, or if the critique collapses into hostile theatre, the tool halts and keeps the original framing visible rather than pretending a cleaner one exists.

pass pass with risk fail no finding insufficient evidence not applicable

Every review returns one of six honest verdicts. The rules forbid the most common dodge: calling something "no finding" when the truth is the evidence was too weak to judge.

7 The spin doctors, in depth

Two adversaries on the same text — and the register they produce together.

Before a high-stakes decision goes public, its wording faces a controlled pair: one voice attacks, one defends. They are never allowed to run apart. The attack alone would be a weapon; the defense alone would be blind to what it is defending against.

pro-spin · the attack

Models the hostile reading

Takes the frozen text and builds the strongest distortion a real adversary would apply — what gets clipped, what gets read out of context, where urgency is inflated, where a quote can be turned.

it may

  • model clipping and context collapse
  • model author-versus-narrator confusion
  • inflate urgency where the text invites it

it may not

  • fabricate motive, biography, or affiliation
  • turn contradiction into proven hidden intent
  • invent an attack the source could not trigger
anti-spin · the defense

Hardens without hiding

Answers the strongest hostile frame and rewrites the wording to reduce the attack surface — while refusing to bury any real tradeoff to do it. Truth-preserving is the binding constraint.

it must

  • preserve what is true under pressure
  • answer the strongest plausible hostile frame
  • separate evidenced consequence from inferred motive
  • cut attack surface without hiding a tradeoff

The attack lane may never ship alone. A packet that carries pro-spin without its anti-spin companion is, by rule, blocked from publication.

where the value compounds

Overlay the two, and a third artifact appears.

Run as a pair, the spin doctors give an attack and a defense. Fused into one document, they produce something neither writes alone: a line-by-line wording-risk register. For every vulnerable passage, three columns — the attack it invites, the hardening that answers it, and the residual risk that honesty would not let us defend away.

Illustrative — how the register reads in practice.

Passage at riskAttack it invitesHardened wordingResidual risk
"deployed in production" clipped to imply a live public service "a working local runtime, not a public service" none — claim now matches the build
"100% consensus" read as proof the verdict is certain "full agreement on this run; logs attached" a critic may still question the sample
"signed by the council" read as legal sign-off "trace-bound to the council's vote record" "signed" carries weight we must scope in person

The residual column is the deliverable a minister actually needs: the honest list of what will still be used against the text after every safe fix is made. The pair finds the attacks. The overlay tells you which ones you cannot wording your way out of — so you decide them in the open, before they decide themselves in public.

8 What this offers policy

Auditability that does not depend on trusting the operator.

01

Evidence over assurance

Every decision ships with its triage tier, its votes, its recorded dissent, and its certification results — not a vendor's word that review happened.

02

Plurality without groupthink

Seats are modelled on genuinely different experts across economics, law, ethics, philosophy, and international politics — and certified to stay distinct under pressure.

03

Risk-scaled cost

Low-stakes decisions move fast on a small panel. Only high-stakes cases convene the full thirty-plus-lane council. Governance scales with consequence.

04

A boundary you can inspect

The line between building and certifying is the one claim a regulator can check directly. The method is designed so that line is always visible.