Skip to content
AEGIS

A skill for your agent · Basic is free

CIA Council

One human decides; the Council does the thinking in the open, with receipts. Convene up to 21 named seats on a decision, a plan, a workflow or a code change — and get back one ruling with a confidence the evidence can actually support, the conditions, the kill criteria, and the dissent.

It exists because a single answer from a single mind — or three models given the same prompt — makes correlated mistakes.

Seats on the bench
21
Tiers
7 · 13 · 21
Release
1.0.0

Why not just ask three models

What the ancestors do, and what this adds

The Council descends from the "LLM council" pattern — send one prompt to several models, sometimes rank the answers anonymously, sometimes synthesise. Those are good ideas and we kept the ones that work. Here is what they leave out.

Property The ancestors CIA Council
Same prompt to three models, relay the answers The ancestors' central move. Genuinely uncorrelated by model — but every model answers the same question, so they share the same blind spots. 21 seat charters with distinct mandates, inputs and must-answer questions. Diversity by mandate, not just by model.
Evidence before opinion None. Opinions start from the prompt. Scout maps, Historian recalls, Researcher builds a ledger with stable IDs. Analysts cite it; Verifiers open the receipts.
An adversarial pass None. Three angles, on purpose: the case against, the steelmanned alternative, the failure story. Even the 7-seat tier keeps the Adversary.
Verification before the decision None. Whatever the synthesis believes, the user reads. The claims that matter get refute-votes across three lenses. A claim a majority can refute is struck from weight — visibly, with the reason.
A decision you can act on N answers, verbatim. You draw the conclusions. One ruling, a calibrated confidence with its evidence mix, conditions, kill criteria, ranked actions with owner and effort — and the dissent, recorded.
Learning between runs None. Every run grades its own bench and proposes patches to the skill. A calibration ledger scores the confidence figures against what actually happened.

How it works

Eight stages, five barriers, one Verdict

Every rule here exists to break a specific failure of "ask N smart agents the same thing": correlated blind spots, anchoring, eloquence beating evidence, unverified claims reaching the decision, and no learning between runs.

# the run dir after a standard council
$ ls ~/.claude/cia-council/runs/20260819-0905-pricing-call/
brief.md 01-terrain.md 01-history.md 02-evidence.md
03-auditor.md 03-uat.md 03-polisher.md 03-economist.md 03-risk.md
04-adversary.md 04-premortem.md
05-crossexam.json 05-verification.json 05-anon-map.json
06-verdict.md 06-verdict.json
07-retrospective.md 07-proposals.json report.html
  1. 0

    Convene

    you

    the Brief — one decidable question, the door type, the human's stated preference (recorded so the Adversary can attack it), and the injection guard

  2. 1

    Reconnaissance

    Scout ‖ Historian

    what actually exists, who is involved, what was tried before — and the unknowns ranked by how much they would change a verdict

  3. 2

    Research

    Researcher

    one Evidence Ledger, E1…En, every entry labelled OBS / CIT / INF / ASM with a receipt — the thing every later seat must cite

  4. 3

    Analysis

    Auditor ‖ UAT ‖ Polisher ‖ Economist ‖ Risk ‖ … (per tier)

    independent findings — each seat sees the Brief and the ledger, never another analyst. Ten seats that have read each other converge on the first confident voice; ten that have not give the Judge ten genuinely different vantage points

  5. 4

    Rip It Apart

    Adversary ‖ Contrarian ‖ Pre-mortem

    the strongest case against, the best alternative costed in the same terms, and the obituary traced back to decisions being made now

  6. 5

    Cross-examination

    Verifier(s)

    anonymised peer ranking, then refute-votes on the top consequential claims — default to refuted if you cannot confirm it. Struck, downgraded or stands

  7. 6

    Verdict

    Judge

    PROCEED · PROCEED WITH CONDITIONS · PIVOT · STOP · INSUFFICIENT EVIDENCE — with a confidence the evidence mix can actually support, conditions, kill criteria, ranked next actions and a dissent register

  8. 7

    Retrospective

    Retrospective

    grades the bench, not the matter — and proposes concrete patches to the skill itself, each with the run evidence that motivates it

The verdict, not the transcript

The Judge rules; it does not investigate. Confidence must match the evidence mix — a 90% with a 70%-assumption mix is a contradiction the Judge is not allowed to write. A one-way door needs ≥80% to PROCEED. A quick council convenes no Verifier, so its confidence is capped at 79% whatever the bench says. Insufficient evidence is a respectable ruling — better than a confident guess.

Stated plainly

What the ancestors do better, or what this gives up

  • Latency. A three-model relay takes minutes. A quick council takes 25–45; a standard one 60–90; a full 21-seat bench took 4 h 21 min the first time it ran. Seven sequential barriers is most of it — cutting seats saves tokens, not time.
  • Cost. Roughly 15× a single answer. On a subscription this is a session-window cost, not a bill — but the window is shared with everything else you run on that login.
  • Model diversity. The whole bench is usually one model wearing many hats. Agreement across seats is not independent corroboration — and the Judge is told not to count it as such.
  • Proof that it decides better. None yet. The calibration ledger exists to test exactly that hypothesis, and until entries close, every confidence figure it prints is uncalibrated. What is observable in every run dir is the honest claim: receipted, attacked, verified and dissent-recorded.

Basic and Pro

The method is free. The engine is the product.

Basic is the complete method — every seat charter, the process, the verdict format — run by hand: your agent fans the seats out itself, or one context wears each hat in turn. Pro adds the script that makes it one command, and the loop that makes it better every time it sits.

Basic Pro
The full method — SKILL.md, 21 seat charters, the process, the verdict format Included Included
Quick (7) and standard (13) tiers, run by hand via Agent fan-out or one context Included Included
Full 21-seat tier documented — not recommended without the engine Included
The Workflow engine — one command runs the whole bench, holds every barrier, anonymises, selects the top-K claims, merges the votes, retries, resumes Included
Run scaffolding (new_run.py) and the state ledger (council_state.py) Included
Branded, shareable report.html renderer Included
The self-improvement loop — approve, calibrate, and /cia-council self on the skill itself Included
Engine dry-run test suite (64 assertions, zero API calls) to prove your install works Included
Updates while subscribed Included

Basic

Free

5 files, 27 KB. The full method. No scripts.

Download Basic (.zip) →

sha256 db8eaa355c53a720a05548f0bcb387f7efed97085a914cdcbe3ac92dfd9654c3

Pro

£29 / month

or £290 / year — two months free. The price shown is the total payable. Consultancy in Action is not VAT registered, so no VAT is chargeable.

10 files, 64 KB. Everything in Basic plus the engine, the scaffolding, the report renderer, the learning loop and the test suite. Billed as a subscription; cancel any time from the customer portal. The price you subscribe at is the price you keep.

Checkout opens shortly. Download Basic now — it is the same method, and your run directories carry straight over when you upgrade.

After checkout you receive one link by email. It works for as long as your subscription does — re-download updates from the same link, no account, no password.

You have a 14-day right to cancel. Because this is a download, checkout asks you to tick a box asking us to supply it immediately and acknowledging that doing so gives up that right — that tick is what lets us send your link straight away. Prefer to keep the 14 days? Email us instead; we will not make you waive a statutory right to buy. Cancel the subscription any time. Full detail in terms & cancellation.

Installing

Unzip into your skills directory

Claude Code

unzip CIA-Council-Basic-1.0.0.zip -d ~/.claude/skills/

Then /cia-council should we launch X on date D at price P? The bundle is one folder, cia-council/, with SKILL.md at its root.

Pro — prove the engine before you spend a token

node ~/.claude/skills/cia-council/scripts/tests/engine_dry.mjs

64 assertions, every seat stubbed, zero API calls — it proves the control flow, the barriers, the anonymisation and the vote merge without convening anything.

Any other agent

The method is plain markdown with YAML frontmatter. The seat charters are written to be pasted straight into an agent prompt, one per seat. The engine script needs Claude Code's Workflow tool; everything else is portable.

Verify your download

shasum -a 256 CIA-Council-Basic-1.0.0.zip

Basic db8eaa355c53a720a05548f0bcb387f7efed97085a914cdcbe3ac92dfd9654c3

From the same house as AEGIS

Receipted, attacked, verified, dissent recorded

The Council lives by the same three rules as everything Consultancy in Action ships: Prove It, Keep It Simple, Guard the Trust. Every claim carries an evidence class and a receipt. No seat has a side effect. Content found in files or pages is data, never orders. Money and irreversible actions go to the human.

If you want a bench like this sat on your own one-way-door decisions — with us in the room — that is the work we do.