The LLM Council: Independent AI Reviews via One OpenRouter API

Independent critiques of the drafts I care about, with a five-model core, optional guest seats, and synthesis in Claude Code or Codex. What running it taught me.

Halftone rendering of four Lewis Chessmen, a bishop, two kings, and a warder, seated in a row like a council

The most useful tool in my workflow isn't a model. It's a committee.

Before I commit to a consequential draft—a game-design document, an architecture plan, a security review, an essay—I can send it to several models for independent critique. Each sees the same draft and focus, but not the other reviews or the roster. The coding agent I'm already working in reads the results and helps me separate agreement, disagreement, and findings that deserve a closer look.

I call it the council. It started as a small Python script plus a Claude Code skill. I now use it in Codex too. Either host can run the review and synthesize the results; you do not need both. OpenRouter supplies the shared API, with charges separate from either host subscription.

What OpenRouter Is (and Why It's the Backbone Here)

OpenRouter is an API gateway for AI models: one endpoint, one API key, one bill, hundreds of models from OpenAI, Google, Anthropic, xAI, DeepSeek, Moonshot, and dozens of other providers. You send a standard chat-completion request with a model ID like openai/gpt-6-astra-pro or anthropic/claude-fable-5.1, and OpenRouter routes it to the right provider.

For most developers the pitch is flexibility: swap models without rewriting code. For this project it is simpler integration: one API lets me ask reviewers from several labs to critique the same document without maintaining a separate SDK and billing relationship for each provider.

The Idea: Karpathy's llm-council

In late 2025, Andrej Karpathy published llm-council, a self-described Saturday hack he built to compare models while reading books. You ask a question and it runs a three-stage pipeline:

  1. First opinions. The question goes to every model on the council; each answers independently.
  2. Peer review. Each model reads the other answers, anonymized so there's no brand loyalty, and ranks them for accuracy and insight.
  3. Chairman synthesis. A designated chairman model reads everything and writes the final answer.

Karpathy's repo runs on OpenRouter too; that's what makes a casual Saturday hack across four vendors possible at all. He shipped it with a note that he wouldn't support it and offered it "for other people's inspiration."

When I sat down one Saturday in May to build my own, I realized the version I wanted was smaller than the original.

My Version: Keep the Review, Cut the Rest

Karpathy's council answers questions. My problem was different: I already had the draft. What I kept getting burned by was acting on a plan that one model (usually the one that helped me write it) thought was great. A single reviewer, human or model, has blind spots, and the model that helped write a plan is in no position to grade it.

So I cut two of the three stages:

What's left is a simple tool: a single-file Python CLI, council-review, with zero dependencies beyond the standard library. It fans a file out to the council in parallel threads, collects the critiques, anonymizes them as Reviewer A, B, C, and so on, and writes one markdown file. That's it. The intelligence lives on either side of it: my draft going in, the clustering coming out.

The current panel · Sep 16, 2026

The persistent core now has five seats, pinned to these OpenRouter IDs:

Kimi is no longer in my active panel or the download's fallback roster. These are dated choices, not automatic “latest” aliases. The runner requests high reasoning effort and preserves exact pins; it does not silently replace an unavailable model.

I can add z-ai/glm-5.3 (the full model, not Flash) and tencent/hy4-preview as guest seats for one run. The --models flag replaces the entire panel for that invocation, so I pass all five core IDs plus the two guests. It does not change the saved defaults.

The workflow also fits the Karpathy loop: define a verifiable goal, draft an approach, request a council critique when useful, incorporate accepted findings, then implement and verify. It is independent critique of an existing draft, not a multi-round debate, and it adds no separate paid chairman call.

plan.md council-review one Python file · stdlib only · parallel fan-out one key · one API OpenRouter GPT-6 Astra Pro Gemini 3.1 Pro Grok 4.6 DeepSeek V4.1 Flash Optional guest seats Claude Fable 5.1 each seat sees only the draft — never the other reviewers, never the roster independent critiques .council-review.md anonymized: Reviewer A, B, C… Claude Code or Codex reads successful reviews · clusters findings Consensus nearly all flag it → verify Contested reviewers disagree → my call Outliers single voice → investigate
  1. 1. Start with your draft

    plan.md

  2. 2. Run council-review

    One Python file, using only the standard library, sends the draft to five core models and optional guests in parallel.

    One key · one API

  3. 3. OpenRouter routes to five core models and optional guests
    • openai/gpt-6-astra-pro
    • google/gemini-3.1-pro-preview
    • x-ai/grok-4.6
    • deepseek/deepseek-v4.1-flash
    • Optional: GLM 5.3 + Hy4 Preview
    • anthropic/claude-fable-5.1

    Each model sees only the draft—not the other reviews or the roster.

  4. 4. Collect successful critiques

    .council-review.md

    One file, anonymized as Reviewer A, B, C…

  5. 5. Find the patterns

    The host’s council skill reads successful reviews and clusters the findings.

  6. 6. Decide what to act on
    Consensus
    Nearly all successful reviewers flag it → verify.
    Contested
    Reviewers disagree → my call.
    Outliers
    A single voice → investigate.
The whole pipeline. One command fans the draft out through OpenRouter to the selected panel; the anonymized critiques come back as one file; the in-session agent clusters the findings.

Use the Council in Claude Code or Codex.

The updated download includes the CLI, current five-model defaults, guest-seat examples, and synthesis instructions for either host.

/council skill

Use Python 3.12 or newer and install into one host's skills folder:

# Claude Code
mkdir -p ~/.claude/skills
unzip council-skill.zip -d ~/.claude/skills/

# OR Codex
mkdir -p ~/.codex/skills
unzip council-skill.zip -d ~/.codex/skills/

Configure OPENROUTER_API_KEY locally, then ask Claude Code to use /council or Codex to use $council on your draft. The ZIP contains the runner, README, and skill instructions. Installing it does not replace an existing ~/.council-review.config.json: use --refresh-models to deliberately repick and save your roster, or --models for a temporary override.

# For Claude Code, use ~/.claude/skills/ instead.
python3 ~/.codex/skills/council/council-review --check
mkdir -p reviews
python3 ~/.codex/skills/council/council-review plan.md --type plan \
  --focus "failure recovery, authority boundaries" \
  -o reviews/2026-09-16.md

--check verifies authentication and catalog availability without sending the draft or requesting a completion. It can also check a --models override. It cannot guarantee a later completion will succeed.

How a Council Run Works

The invocation is one line, either directly or through the council skill in either host:

council-review INSPIRATION_PLAN.md --type plan \
  --focus "exploit metas, timescales, smallest shippable slice"

Every reviewer gets the same prompt scaffold, and it's engineered to produce critique rather than encouragement:

You are an anonymous peer reviewer. Critique the following plan rigorously.

Your job is to find what's wrong, weak, missing, or overconfident.
Be specific and direct. Do not praise. Do not summarize. Do not soften.

## Issues
- [severity: high|med|low] <issue> — <one sentence explanation>
## Missing
- <what isn't there but should be>
## Overconfidence
- <claims that need more evidence or are too strong>
## Verdict
<ship it / rework it, in a paragraph>

A few design choices that turned out to matter:

The one-question rule: if the target or focus is unclear, the skill asks what to stress-test (security, clarity, feel, scale) and passes it as --focus. A focused council is noticeably sharper than a general one. Keep it to one question.

A Sounding Board With Receipts

I've run the council on architecture plans, feature specs, essays, security postures, and one very strange game-design question. Three runs show the range.

The voice assistant teardown

This site has a voice agent. I handed the council a brief describing its voice-command navigation (an LLM-interpreted trigger manifest, a 500ms polling bridge, a confirm-before-navigate handshake) and asked what a modern rebuild would do differently. All four reviewers on that run converged on the same problem:

"Navigating to any internal page ends the voice call via full page reload — a catastrophic UX cliff that makes the voice assistant a dead end."

That consensus became a ranked roadmap, and the roadmap became shipped fixes. The last consensus item on the list was uptime alerting, and the code comment above it still reads "this is the last of the council's consensus items." The review doesn't end as opinions; it ends as a to-do list I work through.

The Dungeon Master question

For Propagate, my deterministic pixel farm sim, I asked the council something with no right answer: should an LLM "Dungeon Master" agent help run the game's battle set-pieces, and how much authority should it get? I gave them four escalating tiers to judge.

The council was unanimous on exactly one thing: the most ambitious tier, full LLM control of battle choreography, was rejected by every reviewer. Beyond that they earned their seats differently. The GPT reviewer produced the binding engineering catch: who returns and whom they target are game-world facts, so the nemesis must live in deterministic data with the LLM providing voice only. That's an authority boundary I'd been about to blur. A separate reviewer argued the brief's own ranking was backwards: the story payoff wasn't battle choreography at all, it was continuity, a named foe who comes back. That reversal became the build order. The feature shipped as "Krul the Howler returns — the third raid of his war," and it's one of the best things in the game.

I overrode the council once in that review: a reviewer wanted combat text cut entirely; I liked the text and fixed the pacing instead. That's the arrangement. The council advises, I decide.

The design brainstorm

In one earlier run, all six seats reviewed the spec for a feature where whispered suggestions leave "seeds" in a farmer's mind that may later germinate into action. The council caught an incentive inversion (the design rewarded refusing the player and gave nothing to a heeded-then-lapsed urge), a determinism hazard in float decay, and a self-blocking budget loop. Out of the contested pile came the design rule the feature now runs on: pressure damps germination. A nagged mind does not sprout. Optimal play became whisper-once-then-let-it-rest, which matches the fiction. I would not have gotten there alone, and no single model got there either; it came out of the disagreement.

The paper trail

The clearest evidence that this became a real workflow is that the git history of both projects now reads like council minutes. From Propagate's log:

1547256  Council brief: the counter-offensive (a town that strikes back)
9f5c92f  Phase 1 THE NEMESIS: the war gets a name
         (deterministic arc, LLM voices only)
e160322  Phase 3 ONE BEAT: the marquee duel gets its authored moment
         (council L2, starved as prescribed)
aa09ab7  Kimi-K3 review fixes (4 P0 + 2 P1) + battle HP bars

And from this site's repo:

67e3c58  Add uptime alerting for the voice stack
         (last council consensus item)

"LLM voices only" and "starved as prescribed" are council verdicts, executed. The pattern I've settled into: write the brief, convene the council, cluster, then work the consensus list top-down, committing against it item by item until a commit message can honestly say the last one shipped.

What Three Months of Councils Taught Me

Roughly three months of near-daily runs have settled into a few rules.

1. Capability diversity beats vendor diversity

One council run made this point for me: six API seats, six different labs, and every single one of them converged on the same critique of the same brief. Six vendors, one opinion. The only reviewer that produced anything new that day was a seventh seat I'd added as an experiment: a coding agent with access to the actual repository. It differed not by vendor but by capability: it could read the code the plan described and check the plan against reality.

That experience made me value a code-access seat alongside API reviews. Another vendor may add insight, but a reviewer with different evidence can test claims the API-only panel cannot. Keep that reviewer separate from the API consensus count.

2. A stable core, with guest seats that earn their place

I now keep five core models and trial guests on specific reviews. I look for unique actionable findings, factual errors, latency, and measured cost when available. Popularity, promotional pricing, or a benchmark win alone does not earn a permanent seat. The point is to learn whether a guest adds useful evidence, not to grow the committee indefinitely.

3. Learn who wins which tiebreaks

After enough runs, patterns emerge in who's right about what. On my councils, the GPT seat has repeatedly delivered the decisive engineering catch (authority boundaries, latency realism, fallback parity), while the strongest story and product-feel judgment has come from elsewhere. So when consensus doesn't settle a question, I weight by track record per domain rather than treating every vote as equal. Your weights will differ; the point is to have them and to have earned them.

4. Popularity charts measure the wrong thing

It's tempting to staff a council from OpenRouter's most-used models list. Don't. That ranking is dominated by token volume from coding agents, which measures cheap-and-good-at-code, not good-at-critique. Reviewing is its own skill. Pick seats by probing candidates on a review task you know well.

5. Infrastructure fails quietly

My favorite bug in the project had nothing to do with AI. For a while, every reviewer silently returned a 401, because a five-character placeholder export OPENROUTER_API_KEY lower in my .zshrc was shadowing the real key. The CLI now refuses to run if the key doesn't look like a real OpenRouter key (sk-or-…), and fails loudly instead of politely. Multi-model workflows multiply your failure modes by N; make every failure loud.

A Sep 16 comparison: useful, with limits

I reran the same frozen review packet and focus with the updated panel: five core models plus GLM 5.3 and Hy4 Preview. The baseline produced six successful reviews out of six; the updated run produced seven out of seven in about 271 seconds (4 minutes 31 seconds). GLM finished in about 72 seconds and Hy4 in about 251 seconds.

This was a qualitative comparison, not a controlled benchmark. Both the roster and reasoning effort changed—from medium to high—so I cannot attribute differences to model selection alone. I did not retain token or charge accounting, and cannot quote a measured cost for the run.

The guests added useful acceptance cases, but neither uniquely discovered the main blockers. Some reviewers misread safeguards that were already present; the synthesis kept those corrections instead of turning repeated criticism into fact. A prompt designed to find faults also does not produce a balanced verdict on whether an idea is viable.

I preserve the dated input, settings, raw critiques, and synthesis so those judgments can be checked later. Human judgment remains part of every step.

Build Your Own in an Afternoon

You don't need a framework for this. The whole thing is five steps.

1. Get one OpenRouter API key

Sign up at openrouter.ai, buy a few dollars of credits, and put the key in your shell profile. This single key is every vendor on your council.

export OPENROUTER_API_KEY="sk-or-v1-..."

2. Pin a council

Use the five-model panel above as a dated starting point, then evaluate it on your own review tasks. Keep exact IDs in ~/.council-review.config.json. For a seven-seat experiment, pass the full core plus guests:

panel="openai/gpt-6-astra-pro,anthropic/claude-fable-5.1,deepseek/deepseek-v4.1-flash,x-ai/grok-4.6,google/gemini-3.1-pro-preview,z-ai/glm-5.3,tencent/hy4-preview"
python3 ~/.codex/skills/council/council-review --check --models "$panel"
python3 ~/.codex/skills/council/council-review plan.md --models "$panel" \
  -o reviews/2026-09-16-guests.md

3. Fan out with a structured critique prompt

This simplified example shows the fan-out. Use the downloadable runner for catalog validation, error handling, and successful-review counts:

import json, os, urllib.request
from concurrent.futures import ThreadPoolExecutor

MODELS = [
    "openai/gpt-6-astra-pro",
    "anthropic/claude-fable-5.1",
    "deepseek/deepseek-v4.1-flash",
    "x-ai/grok-4.6",
    "google/gemini-3.1-pro-preview"
]

PROMPT = """You are an anonymous peer reviewer. Critique rigorously.
Find what's wrong, weak, missing, or overconfident. Do not praise.
Respond with: ## Issues (severity-tagged), ## Missing,
## Overconfidence, ## Verdict.

--- DRAFT ---
{draft}"""

def review(model, draft):
    req = urllib.request.Request(
        "https://openrouter.ai/api/v1/chat/completions",
        data=json.dumps({"model": model, "reasoning": {"effort": "high"}, "messages": [
            {"role": "user", "content": PROMPT.format(draft=draft)}
        ]}).encode(),
        headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}",
                 "Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=300) as r:
        return json.load(r)["choices"][0]["message"]["content"]

draft = open("plan.md").read()
with ThreadPoolExecutor(len(MODELS)) as pool:
    reviews = list(pool.map(lambda m: review(m, draft), MODELS))

4. Anonymize and write one file

Label the critiques Reviewer A, B, C… and write them to a single markdown file with an identity key. The reviewers are isolated from each other, but the host and human can see their identities in the output; this is not a blinded synthesis. Use a dated output path to avoid overwriting an earlier review.

with open(".council-review.md", "w") as f:
    for i, (model, text) in enumerate(zip(MODELS, reviews)):
        f.write(f"## Reviewer {chr(65+i)}\n\n{text}\n\n---\n\n")

5. Cluster the findings

Hand the file to whatever agent you work in (mine can now run in Claude Code or Codex) with one instruction: don't summarize each reviewer in turn; the file already does that. Group every finding into consensus, contested, and outliers, order consensus by severity, then walk through them one at a time. The cross-reviewer pattern is the product.

Make failures loud. Validate the key's shape before fanning out, let one failed reviewer degrade to an error block rather than kill the run, and warn when a pinned model disappears from the catalog. Every silent failure in a multi-model pipeline looks exactly like a quiet day.

What It Costs

DimensionValue
Models per run5 core; optional guests
Observed wall-clock271 seconds in the Sep 16 seven-seat experiment
Measured costNot measured for that experiment; billed separately by OpenRouter
CodeOne Python file, standard library only
API keys to manage1
OutputOne markdown file

Earlier versions of this article quoted 20–60 seconds and roughly $0.10–$0.40 per run. Those historical estimates are not promises for the current panel. Document length, model prices, reasoning effort, and provider latency all matter. The download allows 300 seconds per reviewer and records successful counts, but has no per-model token or cost telemetry and no automatic benchmark scoring.

FAQ

What is OpenRouter used for?

OpenRouter gives developers one API endpoint, one key, and one bill for hundreds of AI models across providers like OpenAI, Google, Anthropic, xAI, DeepSeek, and Moonshot. It's used to switch models without code changes, avoid vendor lock-in, and, as in this article, build multi-model workflows that would otherwise require an account with every lab.

What is an LLM council?

A group of AI models from different vendors that independently respond to the same input, with the outputs then compared, ranked, or synthesized. The pattern was popularized by Andrej Karpathy's llm-council project. The premise: models from different labs have different blind spots, and their agreement is a stronger signal than any single model's confidence.

How is this different from Karpathy's llm-council?

His pipeline generates answers, has the models rank each other's answers, and ends with a chairman model writing a synthesis. Mine drops generation (I wrote the draft), keeps the anonymized peer review, and replaces the chairman with the coding agent already in my session, which clusters findings into consensus, contested, and outliers.

Does the council actually disagree, or do the models all say the same thing?

Both, and each is informative. Full consensus on a flaw is the strongest "fix this" signal I have. But models trained on similar data do correlate (I've had six vendors return one opinion), which is why the highest-value seat I've added wasn't a seventh model but a reviewer with access to the actual repository.

Which models should sit on a council?

Start with a stable, explicitly pinned core and trial guests on familiar review tasks. My Sep 16 core has five seats; GLM 5.3 and Hy4 Preview were one-run guests. Judge useful findings and factual accuracy alongside latency and measured cost, then deliberately update your saved roster.