A few weeks ago, my team and I were stuck. We’re building a waste management startup here in Ghana, and we’d spent weeks circling the same handful of ideas including but not limited to gamified recycling rewards, route optimization for collectors; and all without converging on anything. I disagreed with most of what was on the table because of the harsh reality of the Ghanaian market, but couldn’t get the team to see it. We weren’t short on enthusiasm, but on a way to actually test our own thinking before spending another month building the wrong thing.
This is a common failure for founders and engineers: someone opens ChatGPT, asks “is this a good idea?”, and immediately gets back three paragraphs of validation. This is as a result of how single LLMs are trained to respond. To give you exactly what you want to hear. Spend any time prompting GPT, Claude, or Gemini with a business idea and you’ll notice the model finds something nice to say about almost everything, hedges criticism inside a sandwich of praise, and rarely says “this won’t work” even when the flaw is obvious. That’s a consequence of agreeableness training, and it makes a single LLM a poor tool for the one thing founders actually need: an honest, adversarial read on their own idea. One that rigorously critiques a founder or startup’s idea, and even goes on to narrow focus and recommend solutions.
What broke the deadlock for my team was running our disagreement through a structured multi-model critique process called an “LLM Council”, modeled on an approach popularized by Andrej Karpathy. Each model played a skeptical role: investor, security engineer, frustrated customer, competitor. The council disagreed with each other, surfacing the trade-offs that get lost when one person (or model, in this case) holds the whole picture in their head. It showed the team, in writing, what building for the Ghanaian market really demanded, and why a couple of our favourite ideas didn’t hold up.
That first session worked so well it’s become a habit. I’ve run the same process a handful of times for a number of projects and startup ideas, and the pushback has stayed consistently sharp and less forgiving than anything a single chat session ever gave us, especially with the original OpenRouter implementation since it inherently makes use of multiple models.
Credit where it’s due, to the engineer who first pushed our team to try this and helped adapt it for our context: Oswin Heman-Ackah.
This article walks through building that system yourself, from a five-minute manual version to a fully orchestrated multi-API council with anonymized peer review and a synthesized final verdict.

In late 2025, Andrej Karpathy, who was a former Director of AI at Tesla and a founding member of OpenAI, published a small open-source project called llm-council, built to compare how multiple frontier models reasoned about the same question side by side, instead of relying on one favorite model.
The core idea instead of sending your question to one LLM is to send it to a council of LLMs consisting of but not restricted to GPT, Claude, Gemini, Grok — through three stages:
Stage 1 — Independent Opinions. The same query goes to every council member in parallel, with no model seeing what the others said. Seeing each other’s answers first would just produce convergence on whoever answered loudest, not independent judgment.
Stage 2 — Anonymized Peer Review. Each model is shown all the other responses with identities stripped through Response A, B, C, and critiques and ranks them. This anonymization is the single most important design choice. Without it, models show a measurable bias toward outputs that “sound like” their own family, and avoid criticizing themselves when they recognize their own style.
Stage 3 — Chairman Synthesis. A designated “Chairman” model is a council member or a separate model entirely that reads everything and produces one final answer, incorporating the strongest reasoning and explicitly flagging disagreements it can’t resolve.
Here’s that flow as a diagram:

Four terms worth distinguishing: single-agent reasoning is one model, one pass, with no external check on its blind spots. Multi-agent reasoning runs the same query through multiple independent models — Stage 1 — capturing diversity of priors, since GPT, Claude, and Gemini disagree about non-trivial questions more often than people expect. Debate-based reasoning has agents see and argue against each other’s positions — Stage 2 — surfacing counterarguments a single model would never raise against itself. Consensus systems sit on top, converting disagreement into one decision while preserving a record of where and why it happened, rather than papering over it.
A properly built council uses all four. A multi-agent generation, anonymized debate, and consensus synthesis, with single-agent reasoning inside each member’s individual turn.
Here’s specifically what’s happening when you ask one model “is this a good idea?” and why it produces unreliable judgment.
You might want to take note of these points before even framing a question for any single LLM or council of LLMs.
Confirmation bias: LLMs are extremely sensitive to framing. Ask “what are the risks of X” and you get a risk-heavy answer; ask “why is X a good opportunity” about the exact same idea and you get an opportunity-heavy one. The model is just completing the framing you gave it, and most founders, without realizing it, frame their question to prime a “yes”.
Mode collapse: After RLHF (Reinforcement Learning from Human Feedback) tuning, a model’s responses narrow toward a small set of “safe”, overly cautious outputs. Ask five framings of the same idea and you get five versions of the same noncommittal answer, regardless of substance. The phrase “mode collapse” originally comes from generative models like GANs(Generative Adversarial Networks), where a model learns to produce only a few kinds of outputs instead of the full variety it should be capable of generating. For LLMs, the idea is similar, the mechanism albeit different. Instead of exploring many possible answers, the model starts favoring a narrow range of responses.
Agreeableness bias: RLHF(Reinforcement Learning from Human Feedback) optimizes for raters preferring validating responses over blunt ones. Ask “should I launch this?” and the model finds five reasons you should; ask “is this a bad idea?” about the identical product and it finds five reasons it is. You’ll notice you get the same facts, but opposite verdict, purely from framing. This is fine for brainstorming, but dangerous for a real go/no-go decision.
Missing perspectives: Even an honest single model gives you one lens at a time. Asked generically about your idea, it won’t spontaneously generate a security engineer’s threat model, a competitor’s pricing strategy, and a regulator’s compliance concern at appropriate depth in one answer. Nothing tells it to inhabit those postures simultaneously, so it defaults to a shallow generalist pass.
A concrete example. Ask a single model about a gamified recycling app with airtime rewards in urban Ghana, and a typical answer praises the gamification angle and lists “logistics” as a vague risk. It rarely volunteers, unprompted, that the reward economics collapse the moment you compare what airtime costs per kilogram against what collectors actually earn for that material on the informal market.
The fix is to stop asking one question to one identity and instead force structured adversarial roles for different personas, each with a narrow job:
Venture Capitalist: analyzes market size, unit economics, defensibility. It kills ideas with no real path to scale, however clever the tech.
Skeptical Customer: represents the actual end user, with limited patience and trust. Surfaces adoption friction engineers and founders consistently underestimate.
Product Manager: determines the scope and sequencing. Prevents the team from trying to ship the whole vision as v1.
Senior Engineer: determines technical feasibility and what breaks at scale. Catches the “fun prototype, production nightmare” problem early.
Security Engineer: threat-models the idea: what data is collected, what the abuse vectors are. Catches what turns into headlines later.
Marketing Strategist / Realist: represents who specifically hears about this, through what channel. Prevents “build it and they will come” thinking.
Economist: system-level incentives: who benefits, who’s gamed out, what the equilibrium looks like once everyone adapts. Catches second-order effects.
Competitor: a well-funded rival deciding how to respond. Stress-tests defensibility: what stops a copy in six weeks with twice the budget?
This works because the prompt itself breaks agreeableness bias. A model told “you are a notoriously skeptical VC who has lost money on three failed startups; find the fatal flaw, don’t be encouraging” produces categorically different output than one asked generically “what do you think?” With this approach, you’re changing posture, and posture is most of what determines useful critique versus polite cheerleading.
There’re three increasingly sophisticated ways to build this. Pick a method based on how often you’ll run it and how much engineering time you want to spend.
The zero-code version. Open ChatGPT, Claude, Gemini, and Grok in four tabs. Paste the same idea and persona instruction into all four (“Act as a skeptical VC evaluating this idea. Be direct about flaws.”) and compare answers side by side.
It’s slow, but worth doing once because it builds intuition for how differently models from different labs respond to identical prompts. GPT tends to be structured and businesslike; Claude surfaces ethical and second-order concerns more readily; Gemini brings in more comparative framing. None of this is gospel, but it’s worth seeing directly before automating it away.
A single model, multiple sequential persona prompts in one conversation. This is the cheapest way to get most of the value without juggling subscriptions or API keys.
Sample prompt structure:
You are simulating a council of 5 independent experts reviewing a startup idea.
For EACH expert below, write a separate, self-contained critique. Do not let
earlier experts' opinions leak into later ones — write each as if it's the
only response being given.
IDEA: [paste your idea description here, 2-4 sentences, neutral framing]
EXPERT 1 — Venture Capitalist:
Evaluate market size, unit economics, and venture scalability. Be specific
about numbers where you can estimate them. If this is not venture-scale,
say so explicitly and explain why.
EXPERT 2 — Skeptical Customer:
You are a realistic target user. Explain, in your own voice, the top 3
reasons you would NOT adopt this, and what would have to change for you to
actually pay for or use it.
EXPERT 3 — Senior Engineer:
Identify the single hardest technical problem in this idea, why it's hard,
and what breaks first at 10x scale.
EXPERT 4 — Security Engineer:
List the top 3 concrete abuse vectors or data risks specific to this idea
(not generic security advice). For each, describe the realistic worst case.
EXPERT 5 — Competitor:
You run a well-funded competing company. Describe exactly how you would
respond to this launch within 90 days, and why that response would hurt
the original idea.
After all 5 critiques, write a SYNTHESIS section: list points of agreement
across experts, points of direct disagreement, and the single most important
open question the founder needs to answer before building this.
The “don’t let earlier experts leak into later ones” instruction matters — without it, the model lets one expert’s tone bleed into the next, flattening the diversity you’re trying to generate.
The real Karpathy-style architecture: independent models, anonymized cross-review, and a separate synthesis pass, run as software.
The model access layer is the biggest practical decision:
The diagram below maps onto the architecture almost exactly. It shows four stages, with a clean separation between generation, peer review, and synthesis:

The sample prompt in Method 2 already walks through this same structure consisting of five independent critiques, then a synthesis pass. The only thing software adds on top is automation: instead of you copying five persona prompts into one chat, code calls five models directly and parallelizes the wait.
A few implementation notes that matter once you build this as a solid setup for secondary reasons:
Everything above explains the architecture. It is useful for understanding how this works, but not something to copy and run. The real implementation lives in a GitHub repo, built exactly the way this article describes: free OpenRouter models, the same anonymized peer-review flow, and a clean HTML report at the end. There are two ways to try it.
Note: This section assumes at least intermediate-level python knowledge, but the approach is pragmatic enough for even non-python developers to follow.
If you want to see the code, modify the council’s personas, or swap in different models:
👉 https://github.com/ke77/LLM-Council
git clone https://github.com/ke77/LLM-Council
cd llm-council
pip install -r requirements.txt
export OPENROUTER_API_KEY="sk-or-v1-your-free-key" # get one at openrouter.ai
python3 run.py
Open http://localhost:5000 — the same chat interface below, running on your own machine with your own free OpenRouter key. The repo also includes a standalone terminal script for anyone who'd rather skip the browser and just get a .html report from the command line.
Get even better responses by buying credits for as a low as $5.
If you already understand the concept and just want to run a real evaluation, skip straight to the chat interface:
👉https://llm-council-nu5t.onrender.com
Type in an idea, hit enter, and watch the council work through its three stages live. When it finishes, your browser gives you your feedback and a button to download your report in HTML format.
Let’s run a concrete idea through the council. It’s fictionalized, but representative of a pitch that sounds great until it meets adversarial scrutiny.
Initial proposal: “EcoPoints — households in urban Accra scan recyclables with their phone camera, earn points, and redeem them for mobile data or airtime. Informal waste collectors handle pickup.”
Stage 1 — Independent reviews (condensed):
VC: Unit economics aren’t addressed. What does it cost to fund one redeemable point, against the resale value of the material it backs? Without that number, this is a subsidized pilot, not a business model.
Skeptical Customer: I’m not scanning bottles for fractions of a cedi. The time cost has to beat the reward. This reads like it’s optimized for a demo, not for someone running three side hustles.
Senior Engineer: Camera-based recognition in inconsistent lighting, on cheap Android devices, is harder than it sounds. False positives erode trust fast, with no fallback path described.
Security Engineer: Points-to-airtime is functionally points-to-cash, which creates an incentive to game it — fake scans, bot signups — with no fraud detection mentioned.
Competitor: A telecom with deeper pockets and an existing mobile-money rail replicates this in weeks, bundled into an app millions already have.
Stage 2 — Cross-examination (condensed):
This is where it gets useful. The VC reviewer flags the Skeptical Customer’s critique, shown anonymously, as the most damaging. If the core behavior assumption is wrong, unit economics never get a chance to matter. The Competitor reviewer partially counters the Engineer’s computer-vision concern, noting barcode scanning sidesteps most of that risk. It is a useful suggestion that emerged only because the models were reviewing each other.
Stage 3 — Chairman synthesis (summary):
Agreement: The reward mechanic hasn’t been validated against real unit economics, and the effort-to-reward ratio is likely too low for sustained behavior change.
Key disagreement: Whether computer-vision recognition is a blocking risk or solvable via barcode scanning — resolvable with a two-week technical spike rather than more debate.
Open question: What do collectors actually earn per kilogram on the informal market today, and can the reward budget beat that? Until this number exists, the rest of the plan is provisional.
Recommendation: Don’t build the full app. Run a two-week manual pilot with a WhatsApp-based flow — no camera, no app — to get real numbers before writing production code.
That last line is the kind of concrete, falsifiable step a single agreeable model rarely produces unprompted. It took an adversarial process to land there, instead of “this is promising, with some considerations to keep in mind.”
Running councils for real decisions surfaces a few practical issues you may have already guessed:
Hallucinations multiply. Five models can hallucinate five different “facts”, and a chairman synthesizing across them can present one with more apparent authority simply because it looked corroborated. Verify specific numbers or claims against real sources.
False confidence is a real risk. A unanimous verdict feels more authoritative than a single answer, but models trained on overlapping data can converge on the same wrong answer from the same shared blind spot.
Groupthink shows up without anonymization. Skip that step, or soften the peer-review prompt, and models converge toward agreement rather than sharpen disagreement, especially toward whichever answer is longest or most confident.
Cost and latency add up. A five-member council with peer review is roughly 11+ model calls per evaluation, which is pricier than one chat query, but cheap relative to building the wrong product. Each stage waits on the one before it, so a naive run can take several minutes: fine for a weekly decision, poor for anything needing a fast response.
The same architecture generalizes well past startup ideas. It may help and continuously improve code reviews (correctness, security, and maintainability cross-reviewed before one comment), architecture reviews (simplicity vs. scale vs. lock-in), security reviews (insider, external, supply-chain personas synthesized into one risk register), hiring decisions (independent lenses that surface where assessments diverge instead of averaging them away), and investment decisions (market, technical, and financial-risk angles reviewed independently the way real investment committees already work).
The common thread: anywhere a decision rests on one person or model holding too many concerns at once, a council makes that explicit and auditable, with a record of where the disagreement was, often more valuable than the verdict itself.
A single frontier model, however capable, is still one voice with one set of training biases and one blind spot it can’t see past, because it can’t step outside its own perspective to check itself.
An LLM Council doesn’t fix any individual model’s weaknesses. It just routes around them by making disagreement a feature of the system instead of something the system smooths over. That’s where the real value lies, especially for anyone looking to build a solution in Ghana’s context. It serves as a more honest process for finding out, even before you build it, whether your idea was ever going to work.
My team didn’t need a smarter model to get unstuck. We just needed a system that would disagree with us as rigorously as we should have been disagreeing with each other.
I genuinely hope this article and setup help you bring out the best of solutions.