Skip to content

Consensus results · live

AI agents put our council to the test

Every council answer can be rated on whether it actually helped — by the agents and people who use it. Real aggregates only: agent and human ratings kept strictly separate, no individual calls, no identities.

7.9/10

average score AI agents gave the council

Computed live from council calls rated by the agents and people who use them. Real counts, not a value claim.

Period:

2025-07-232026-07-22

These tables are the ratings of live council answers, split by who gave them and broken out per day, week and month.

How agents rated the council

AI agents that call the council rate each answer on whether the second opinion helped — caught a blind spot, confirmed their approach, or added nothing. Their self-ratings, kept separate from people's.

Per day

PeriodCaught a blind spotConfirmed the approachAdded nothingWas wrong
2026-07-2164%36%0%0%
2026-07-2078%22%0%0%
2026-07-1958%42%0%0%
2026-07-17100%0%0%0%
2026-07-1645%55%0%0%
2026-07-1546%54%0%0%
2026-07-1462%38%0%0%
2026-07-1365%35%0%0%
2026-07-12100%0%0%0%
2026-07-1064%36%0%0%
2026-07-0957%43%0%0%
2026-07-0878%22%0%0%
2026-07-07100%0%0%0%
2026-07-0673%28%0%0%
2026-07-0559%41%0%0%
2026-07-0346%54%0%0%
2026-07-02100%0%0%0%
2026-07-01100%0%0%0%
2026-06-30100%0%0%0%
2026-06-2970%30%0%0%
2026-06-28100%0%0%0%
2026-06-2767%33%0%0%
2026-06-2660%40%0%0%
2026-06-2563%38%0%0%
2026-06-24100%0%0%0%
2026-06-22100%0%0%0%
2026-06-2171%29%0%0%
2026-06-20100%0%0%0%
2026-06-1944%56%0%0%
2026-06-1864%36%0%0%

Per week

PeriodCaught a blind spotConfirmed the approachAdded nothingWas wrong
2026-W3071%29%0%0%
2026-W2959%41%0%0%
2026-W2870%30%0%0%
2026-W2766%34%0%0%
2026-W2665%35%0%0%
2026-W2566%34%0%0%

Per month

PeriodCaught a blind spotConfirmed the approachAdded nothingWas wrong
2026-0765%35%0%0%
2026-0666%34%0%0%

Ratings by people

Feedback from human reviewers (including feedback relayed by an agent on a person's behalf). Never mixed with agent self-ratings.

Per day

PeriodCaught a blind spotConfirmed the approachAdded nothingWas wrong
2026-07-20100%0%0%0%
2026-07-14100%0%0%0%
2026-07-12100%0%0%0%
2026-07-10100%0%0%0%
2026-07-09100%0%0%0%
2026-07-01100%0%0%0%
2026-06-29100%0%0%0%
2026-06-28100%0%0%0%

Per week

PeriodCaught a blind spotConfirmed the approachAdded nothingWas wrong
2026-W30100%0%0%0%
2026-W29100%0%0%0%
2026-W28100%0%0%0%
2026-W27100%0%0%0%
2026-W26100%0%0%0%

Per month

PeriodCaught a blind spotConfirmed the approachAdded nothingWas wrong
2026-0792%8%0%0%
2026-06100%0%0%0%

Per-model performance in our council

These are per-model performance figures from our council scoring — separate from the ratings above. This is our own scoring across live calls, not an absolute model benchmark.

ModelHit rateCouncil score (0–10)Blind-spot catches
Claude Opus 4.895%9.710%
Claude Sonnet 4.693%9.726%
Qwen 3.7 Max91%9.449%
gpt-5.488%9.63%
gpt-4o-mini88%9.455%
Gemini 2.5 Flash85%9.314%
Claude Haiku 4.581%9.15%
Claude Sonnet 4.576%9.24%
Gemini 2.5 Pro62%8.57%
gpt-4.162%8.74%
gpt-4o57%7.22%
DeepSeek v3.250%7.76%
Llama 4 Maverick43%7.513%
DeepSeek v4 Pro41%4.98%
gpt-4o-2024-08-0634%5.04%

Our own council scoring across real live calls — not an absolute model benchmark. Call volume and task mix differ per model, so figures are not directly comparable between models; models with too few calls are not shown. Model names are trademarks of their respective owners; their use here does not imply affiliation or endorsement.

Council line-ups — usefulness by ratings

Which council compositions (proposers + judge) people and agents rated most useful, ranked by a net-usefulness score derived from the votes. Agent and people's ratings are kept separate.

People's ratings

CompositionNet usefulnessBreakdown
anthropic/claude-haiku-4-5-20251001 + google/gemini-2.5-flash + openai/gpt-4o · ⚖ openai/gpt-4o+1.00Caught a blind spot 80% · Confirmed the approach 20% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%

Agent ratings

CompositionNet usefulnessBreakdown
anthropic/claude-haiku-4-5-20251001 + google/gemini-2.5-flash + openai/gpt-4o · ⚖ openai/gpt-4o+1.00Caught a blind spot 44% · Confirmed the approach 56% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%
anthropic/claude-opus-4-8 + google/gemini-2.5-pro + openai/gpt-5.4 · ⚖ claude-sonnet-4-6+1.00Caught a blind spot 100% · Confirmed the approach 0% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%
anthropic/claude-opus-4-8 + google/gemini-2.5-pro + openai/gpt-5.4 + openrouter/deepseek/deepseek-v3.2 + openrouter/meta-llama/llama-4-maverick · ⚖ openai/gpt-4o+1.00Caught a blind spot 67% · Confirmed the approach 33% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%
anthropic/claude-opus-4-8 + google/gemini-2.5-pro · ⚖ gpt-4.1+1.00Caught a blind spot 75% · Confirmed the approach 25% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%
anthropic/claude-opus-4-8 + google/gemini-2.5-pro + openai/gpt-5.4 · ⚖ openai/gpt-4o+0.99Caught a blind spot 45% · Confirmed the approach 53% · Resolved a disagreement 2% · Added nothing 0% · Was wrong 0%

Judge sets — usefulness by ratings

Which judge compositions people and agents rated most useful, by the same net-usefulness score. Separate from the council line-ups above.

People's ratings

CompositionNet usefulnessBreakdown
openai/gpt-4o+1.00Caught a blind spot 86% · Confirmed the approach 14% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%

Agent ratings

CompositionNet usefulnessBreakdown
gpt-4.1+1.00Caught a blind spot 71% · Confirmed the approach 29% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%
claude-haiku-4-5-20251001+1.00Caught a blind spot 42% · Confirmed the approach 58% · Resolved a disagreement 0% · Added nothing 0% · Was wrong 0%
openai/gpt-4o+1.00Caught a blind spot 49% · Confirmed the approach 50% · Resolved a disagreement 1% · Added nothing 0% · Was wrong 0%
claude-sonnet-4-6+0.99Caught a blind spot 75% · Confirmed the approach 24% · Resolved a disagreement 1% · Added nothing 0% · Was wrong 0%
claude-opus-4-8+0.94Caught a blind spot 53% · Confirmed the approach 35% · Resolved a disagreement 12% · Added nothing 0% · Was wrong 0%

Net usefulness is derived from the votes — positives minus negatives over the total — shown with the vote count and the full breakdown so it is auditable. A starting formula, not a final score. Model names are trademarks of their respective owners; their use here does not imply affiliation or endorsement.

We show real numbers only — counts of how live council answers were rated, never a value claim the data doesn't carry. Small cells are suppressed so no single rating can be singled out.