The most useful tool in my workflow isn't a model. It's a committee.
Before I commit to a consequential draft—a game-design document, an architecture plan, a security review, an essay—I can send it to several models for independent critique. Each sees the same draft and focus, but not the other reviews or the roster. The coding agent I'm already working in reads the results and helps me separate agreement, disagreement, and findings that deserve a closer look.
I call it the council. It started as a small Python script plus a Claude Code skill. I now use it in Codex too. Either host can run the review and synthesize the results; you do not need both. OpenRouter supplies the shared API, with charges separate from either host subscription.
What OpenRouter Is (and Why It's the Backbone Here)
OpenRouter is an API gateway for AI models: one endpoint, one API key, one bill, hundreds of models from OpenAI, Google, Anthropic, xAI, DeepSeek, Moonshot, and dozens of other providers. You send a standard chat-completion request with a model ID like openai/gpt-6-astra-pro or anthropic/claude-fable-5.1, and OpenRouter routes it to the right provider.
For most developers the pitch is flexibility: swap models without rewriting code. For this project it is simpler integration: one API lets me ask reviewers from several labs to critique the same document without maintaining a separate SDK and billing relationship for each provider.
The Idea: Karpathy's llm-council
In late 2025, Andrej Karpathy published llm-council, a self-described Saturday hack he built to compare models while reading books. You ask a question and it runs a three-stage pipeline:
- First opinions. The question goes to every model on the council; each answers independently.
- Peer review. Each model reads the other answers, anonymized so there's no brand loyalty, and ranks them for accuracy and insight.
- Chairman synthesis. A designated chairman model reads everything and writes the final answer.
Karpathy's repo runs on OpenRouter too; that's what makes a casual Saturday hack across four vendors possible at all. He shipped it with a note that he wouldn't support it and offered it "for other people's inspiration."
When I sat down one Saturday in May to build my own, I realized the version I wanted was smaller than the original.
My Version: Keep the Review, Cut the Rest
Karpathy's council answers questions. My problem was different: I already had the draft. What I kept getting burned by was acting on a plan that one model (usually the one that helped me write it) thought was great. A single reviewer, human or model, has blind spots, and the model that helped write a plan is in no position to grade it.
So I cut two of the three stages:
- Stage 1 (generation): gone. I already wrote the draft.
- Stage 2 (anonymized peer review): kept. This is the part that adds signal. Independent critics who can't see each other, can't see who else is on the panel, and are explicitly told not to praise.
- Stage 3 (chairman): replaced. Instead of a separate synthesis model, the coding agent already in my session does the clustering. It has the context of what I'm building; a chairman doesn't.
What's left is a simple tool: a single-file Python CLI, council-review, with zero dependencies beyond the standard library. It fans a file out to the council in parallel threads, collects the critiques, anonymizes them as Reviewer A, B, C, and so on, and writes one markdown file. That's it. The intelligence lives on either side of it: my draft going in, the clustering coming out.
The current panel · Sep 16, 2026
The persistent core now has five seats, pinned to these OpenRouter IDs:
openai/gpt-6-astra-proanthropic/claude-fable-5.1deepseek/deepseek-v4.1-flashx-ai/grok-4.6google/gemini-3.1-pro-preview
Kimi is no longer in my active panel or the download's fallback roster. These are dated choices, not automatic “latest” aliases. The runner requests high reasoning effort and preserves exact pins; it does not silently replace an unavailable model.
I can add z-ai/glm-5.3 (the full model, not Flash) and tencent/hy4-preview as guest seats for one run. The --models flag replaces the entire panel for that invocation, so I pass all five core IDs plus the two guests. It does not change the saved defaults.
The workflow also fits the Karpathy loop: define a verifiable goal, draft an approach, request a council critique when useful, incorporate accepted findings, then implement and verify. It is independent critique of an existing draft, not a multi-round debate, and it adds no separate paid chairman call.
- 1. Start with your draft
plan.md
- 2. Run council-review
One Python file, using only the standard library, sends the draft to five core models and optional guests in parallel.
One key · one API
- 3. OpenRouter routes to five core models and optional guests
- openai/gpt-6-astra-pro
- google/gemini-3.1-pro-preview
- x-ai/grok-4.6
- deepseek/deepseek-v4.1-flash
- Optional: GLM 5.3 + Hy4 Preview
- anthropic/claude-fable-5.1
Each model sees only the draft—not the other reviews or the roster.
- 4. Collect successful critiques
.council-review.md
One file, anonymized as Reviewer A, B, C…
- 5. Find the patterns
The host’s council skill reads successful reviews and clusters the findings.
- 6. Decide what to act on
- Consensus
- Nearly all successful reviewers flag it → verify.
- Contested
- Reviewers disagree → my call.
- Outliers
- A single voice → investigate.
Use the Council in Claude Code or Codex.
The updated download includes the CLI, current five-model defaults, guest-seat examples, and synthesis instructions for either host.
Use Python 3.12 or newer and install into one host's skills folder:
# Claude Code
mkdir -p ~/.claude/skills
unzip council-skill.zip -d ~/.claude/skills/
# OR Codex
mkdir -p ~/.codex/skills
unzip council-skill.zip -d ~/.codex/skills/
Configure OPENROUTER_API_KEY locally, then ask Claude Code to use /council or Codex to use $council on your draft. The ZIP contains the runner, README, and skill instructions. Installing it does not replace an existing ~/.council-review.config.json: use --refresh-models to deliberately repick and save your roster, or --models for a temporary override.
# For Claude Code, use ~/.claude/skills/ instead.
python3 ~/.codex/skills/council/council-review --check
mkdir -p reviews
python3 ~/.codex/skills/council/council-review plan.md --type plan \
--focus "failure recovery, authority boundaries" \
-o reviews/2026-09-16.md
--check verifies authentication and catalog availability without sending the draft or requesting a completion. It can also check a --models override. It cannot guarantee a later completion will succeed.
How a Council Run Works
The invocation is one line, either directly or through the council skill in either host:
council-review INSPIRATION_PLAN.md --type plan \
--focus "exploit metas, timescales, smallest shippable slice"
Every reviewer gets the same prompt scaffold, and it's engineered to produce critique rather than encouragement:
You are an anonymous peer reviewer. Critique the following plan rigorously.
Your job is to find what's wrong, weak, missing, or overconfident.
Be specific and direct. Do not praise. Do not summarize. Do not soften.
## Issues
- [severity: high|med|low] <issue> — <one sentence explanation>
## Missing
- <what isn't there but should be>
## Overconfidence
- <claims that need more evidence or are too strong>
## Verdict
<ship it / rework it, in a paragraph>
A few design choices that turned out to matter:
- Anonymity during, transparency after. Each model sees only the draft, not the other reviews or the roster. The output file reveals the identity key (Reviewer A →
anthropic/claude-fable-5.1, and so on) so I can weigh voices after the fact, but no reviewer can play to the room. - Pinned models, checked before review. Saved configuration overrides the five bundled defaults. Catalog checks cover temporary panels too. Missing seats are reported and omitted; a catalog failure or no available seats stops the run. An explicit
--skip-freshness-checkbypasses the check. - Count real replies. Empty replies and error blocks are failures, not votes. The updated download reports successful counts and returns a nonzero exit status if every reviewer fails. It redacts the configured API key from completion error messages.
- Synthesize by issue. Consensus means all or all but one successful reviewer flags an issue, with at least three successful reviewers. Give counts. Distinguish partial agreement from explicit disagreement: silence is not opposition. An outlier can still identify the most important defect.
- Verify before acting. Check consequential claims against code or primary sources. Repetition does not make a claim true. The host synthesizes; I decide which findings to accept. Review is not permission to implement or ship.
--focus. A focused council is noticeably sharper than a general one. Keep it to one question.
A Sounding Board With Receipts
I've run the council on architecture plans, feature specs, essays, security postures, and one very strange game-design question. Three runs show the range.
The voice assistant teardown
This site has a voice agent. I handed the council a brief describing its voice-command navigation (an LLM-interpreted trigger manifest, a 500ms polling bridge, a confirm-before-navigate handshake) and asked what a modern rebuild would do differently. All four reviewers on that run converged on the same problem:
"Navigating to any internal page ends the voice call via full page reload — a catastrophic UX cliff that makes the voice assistant a dead end."
That consensus became a ranked roadmap, and the roadmap became shipped fixes. The last consensus item on the list was uptime alerting, and the code comment above it still reads "this is the last of the council's consensus items." The review doesn't end as opinions; it ends as a to-do list I work through.
The Dungeon Master question
For Propagate, my deterministic pixel farm sim, I asked the council something with no right answer: should an LLM "Dungeon Master" agent help run the game's battle set-pieces, and how much authority should it get? I gave them four escalating tiers to judge.
The council was unanimous on exactly one thing: the most ambitious tier, full LLM control of battle choreography, was rejected by every reviewer. Beyond that they earned their seats differently. The GPT reviewer produced the binding engineering catch: who returns and whom they target are game-world facts, so the nemesis must live in deterministic data with the LLM providing voice only. That's an authority boundary I'd been about to blur. A separate reviewer argued the brief's own ranking was backwards: the story payoff wasn't battle choreography at all, it was continuity, a named foe who comes back. That reversal became the build order. The feature shipped as "Krul the Howler returns — the third raid of his war," and it's one of the best things in the game.
I overrode the council once in that review: a reviewer wanted combat text cut entirely; I liked the text and fixed the pacing instead. That's the arrangement. The council advises, I decide.
The design brainstorm
In one earlier run, all six seats reviewed the spec for a feature where whispered suggestions leave "seeds" in a farmer's mind that may later germinate into action. The council caught an incentive inversion (the design rewarded refusing the player and gave nothing to a heeded-then-lapsed urge), a determinism hazard in float decay, and a self-blocking budget loop. Out of the contested pile came the design rule the feature now runs on: pressure damps germination. A nagged mind does not sprout. Optimal play became whisper-once-then-let-it-rest, which matches the fiction. I would not have gotten there alone, and no single model got there either; it came out of the disagreement.
The paper trail
The clearest evidence that this became a real workflow is that the git history of both projects now reads like council minutes. From Propagate's log:
1547256 Council brief: the counter-offensive (a town that strikes back)
9f5c92f Phase 1 THE NEMESIS: the war gets a name
(deterministic arc, LLM voices only)
e160322 Phase 3 ONE BEAT: the marquee duel gets its authored moment
(council L2, starved as prescribed)
aa09ab7 Kimi-K3 review fixes (4 P0 + 2 P1) + battle HP bars
And from this site's repo:
67e3c58 Add uptime alerting for the voice stack
(last council consensus item)
"LLM voices only" and "starved as prescribed" are council verdicts, executed. The pattern I've settled into: write the brief, convene the council, cluster, then work the consensus list top-down, committing against it item by item until a commit message can honestly say the last one shipped.
What Three Months of Councils Taught Me
Roughly three months of near-daily runs have settled into a few rules.
1. Capability diversity beats vendor diversity
One council run made this point for me: six API seats, six different labs, and every single one of them converged on the same critique of the same brief. Six vendors, one opinion. The only reviewer that produced anything new that day was a seventh seat I'd added as an experiment: a coding agent with access to the actual repository. It differed not by vendor but by capability: it could read the code the plan described and check the plan against reality.
That experience made me value a code-access seat alongside API reviews. Another vendor may add insight, but a reviewer with different evidence can test claims the API-only panel cannot. Keep that reviewer separate from the API consensus count.
2. A stable core, with guest seats that earn their place
I now keep five core models and trial guests on specific reviews. I look for unique actionable findings, factual errors, latency, and measured cost when available. Popularity, promotional pricing, or a benchmark win alone does not earn a permanent seat. The point is to learn whether a guest adds useful evidence, not to grow the committee indefinitely.
3. Learn who wins which tiebreaks
After enough runs, patterns emerge in who's right about what. On my councils, the GPT seat has repeatedly delivered the decisive engineering catch (authority boundaries, latency realism, fallback parity), while the strongest story and product-feel judgment has come from elsewhere. So when consensus doesn't settle a question, I weight by track record per domain rather than treating every vote as equal. Your weights will differ; the point is to have them and to have earned them.
4. Popularity charts measure the wrong thing
It's tempting to staff a council from OpenRouter's most-used models list. Don't. That ranking is dominated by token volume from coding agents, which measures cheap-and-good-at-code, not good-at-critique. Reviewing is its own skill. Pick seats by probing candidates on a review task you know well.
5. Infrastructure fails quietly
My favorite bug in the project had nothing to do with AI. For a while, every reviewer silently returned a 401, because a five-character placeholder export OPENROUTER_API_KEY lower in my .zshrc was shadowing the real key. The CLI now refuses to run if the key doesn't look like a real OpenRouter key (sk-or-…), and fails loudly instead of politely. Multi-model workflows multiply your failure modes by N; make every failure loud.
A Sep 16 comparison: useful, with limits
I reran the same frozen review packet and focus with the updated panel: five core models plus GLM 5.3 and Hy4 Preview. The baseline produced six successful reviews out of six; the updated run produced seven out of seven in about 271 seconds (4 minutes 31 seconds). GLM finished in about 72 seconds and Hy4 in about 251 seconds.
This was a qualitative comparison, not a controlled benchmark. Both the roster and reasoning effort changed—from medium to high—so I cannot attribute differences to model selection alone. I did not retain token or charge accounting, and cannot quote a measured cost for the run.
The guests added useful acceptance cases, but neither uniquely discovered the main blockers. Some reviewers misread safeguards that were already present; the synthesis kept those corrections instead of turning repeated criticism into fact. A prompt designed to find faults also does not produce a balanced verdict on whether an idea is viable.
I preserve the dated input, settings, raw critiques, and synthesis so those judgments can be checked later. Human judgment remains part of every step.
Build Your Own in an Afternoon
You don't need a framework for this. The whole thing is five steps.
1. Get one OpenRouter API key
Sign up at openrouter.ai, buy a few dollars of credits, and put the key in your shell profile. This single key is every vendor on your council.
export OPENROUTER_API_KEY="sk-or-v1-..."
2. Pin a council
Use the five-model panel above as a dated starting point, then evaluate it on your own review tasks. Keep exact IDs in ~/.council-review.config.json. For a seven-seat experiment, pass the full core plus guests:
panel="openai/gpt-6-astra-pro,anthropic/claude-fable-5.1,deepseek/deepseek-v4.1-flash,x-ai/grok-4.6,google/gemini-3.1-pro-preview,z-ai/glm-5.3,tencent/hy4-preview"
python3 ~/.codex/skills/council/council-review --check --models "$panel"
python3 ~/.codex/skills/council/council-review plan.md --models "$panel" \
-o reviews/2026-09-16-guests.md
3. Fan out with a structured critique prompt
This simplified example shows the fan-out. Use the downloadable runner for catalog validation, error handling, and successful-review counts:
import json, os, urllib.request
from concurrent.futures import ThreadPoolExecutor
MODELS = [
"openai/gpt-6-astra-pro",
"anthropic/claude-fable-5.1",
"deepseek/deepseek-v4.1-flash",
"x-ai/grok-4.6",
"google/gemini-3.1-pro-preview"
]
PROMPT = """You are an anonymous peer reviewer. Critique rigorously.
Find what's wrong, weak, missing, or overconfident. Do not praise.
Respond with: ## Issues (severity-tagged), ## Missing,
## Overconfidence, ## Verdict.
--- DRAFT ---
{draft}"""
def review(model, draft):
req = urllib.request.Request(
"https://openrouter.ai/api/v1/chat/completions",
data=json.dumps({"model": model, "reasoning": {"effort": "high"}, "messages": [
{"role": "user", "content": PROMPT.format(draft=draft)}
]}).encode(),
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}",
"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=300) as r:
return json.load(r)["choices"][0]["message"]["content"]
draft = open("plan.md").read()
with ThreadPoolExecutor(len(MODELS)) as pool:
reviews = list(pool.map(lambda m: review(m, draft), MODELS))
4. Anonymize and write one file
Label the critiques Reviewer A, B, C… and write them to a single markdown file with an identity key. The reviewers are isolated from each other, but the host and human can see their identities in the output; this is not a blinded synthesis. Use a dated output path to avoid overwriting an earlier review.
with open(".council-review.md", "w") as f:
for i, (model, text) in enumerate(zip(MODELS, reviews)):
f.write(f"## Reviewer {chr(65+i)}\n\n{text}\n\n---\n\n")
5. Cluster the findings
Hand the file to whatever agent you work in (mine can now run in Claude Code or Codex) with one instruction: don't summarize each reviewer in turn; the file already does that. Group every finding into consensus, contested, and outliers, order consensus by severity, then walk through them one at a time. The cross-reviewer pattern is the product.
What It Costs
| Dimension | Value |
|---|---|
| Models per run | 5 core; optional guests |
| Observed wall-clock | 271 seconds in the Sep 16 seven-seat experiment |
| Measured cost | Not measured for that experiment; billed separately by OpenRouter |
| Code | One Python file, standard library only |
| API keys to manage | 1 |
| Output | One markdown file |
Earlier versions of this article quoted 20–60 seconds and roughly $0.10–$0.40 per run. Those historical estimates are not promises for the current panel. Document length, model prices, reasoning effort, and provider latency all matter. The download allows 300 seconds per reviewer and records successful counts, but has no per-model token or cost telemetry and no automatic benchmark scoring.
FAQ
What is OpenRouter used for?
OpenRouter gives developers one API endpoint, one key, and one bill for hundreds of AI models across providers like OpenAI, Google, Anthropic, xAI, DeepSeek, and Moonshot. It's used to switch models without code changes, avoid vendor lock-in, and, as in this article, build multi-model workflows that would otherwise require an account with every lab.
What is an LLM council?
A group of AI models from different vendors that independently respond to the same input, with the outputs then compared, ranked, or synthesized. The pattern was popularized by Andrej Karpathy's llm-council project. The premise: models from different labs have different blind spots, and their agreement is a stronger signal than any single model's confidence.
How is this different from Karpathy's llm-council?
His pipeline generates answers, has the models rank each other's answers, and ends with a chairman model writing a synthesis. Mine drops generation (I wrote the draft), keeps the anonymized peer review, and replaces the chairman with the coding agent already in my session, which clusters findings into consensus, contested, and outliers.
Does the council actually disagree, or do the models all say the same thing?
Both, and each is informative. Full consensus on a flaw is the strongest "fix this" signal I have. But models trained on similar data do correlate (I've had six vendors return one opinion), which is why the highest-value seat I've added wasn't a seventh model but a reviewer with access to the actual repository.
Which models should sit on a council?
Start with a stable, explicitly pinned core and trial guests on familiar review tasks. My Sep 16 core has five seats; GLM 5.3 and Hy4 Preview were one-run guests. Judge useful findings and factual accuracy alongside latency and measured cost, then deliberately update your saved roster.