Skip to content

Live Evidence

Why one model is not enough

Real data from every council run we process — updated every 15 minutes. No simulations, no cherry-picked examples.

Blind-spot coverage

A blind spot is a real vulnerability or error that one model silently misses while another model in the same council catches it. The chart below shows which models most often provide the unique catch — the finding no other model in the panel flagged.

Model · Unique-catch rate

  • 1Gpt 4o Mini
    100.0%
    54.6%
  • 2Qwen3.7 Max
    89.2%
    48.7%
  • 3Claude Sonnet 4 6
    48.2%
    26.3%
  • 4Gemini 2.5 Flash
    24.9%
    13.6%
  • 5Llama 4 Maverick
    24.0%
    13.1%
  • 6Claude Opus 4 8
    19.0%
    10.4%
  • 7Deepseek V4 Pro
    14.5%
    7.9%
  • 8Gemini 2.5 Pro
    13.2%
    7.2%

Ranked by unique-catch rate. Only models with sufficient data are shown. Rates are percentages of a model's own events — not relative to other models.

Quality scores

Average quality score (0–100) and ok-rate are computed across all judge evaluations where the model acted as a proposer. Ok-rate = fraction of verdicts rated fully correct.

ModelAvg quality (0–100)Ok-rate
Gpt 5.4 Nano 2026 03 1799.8100.0%
Glm 599.7100.0%
Gpt 5.199.7100.0%
Gpt 5.4 2026 03 0599.7100.0%
Gpt 5.2 2025 12 1199.6100.0%
Gpt 5.299.5100.0%
Gpt 5.1 2025 11 1399.5100.0%
Gpt 5.3 Chat Latest99.5100.0%

Reliability

Noise rate = fraction of model responses the council classifier marks as off-topic or low-signal. Error rate = fraction of API calls that returned an error. Both are averages across all qualifying models.

Avg noise rate

1.98%

Share of responses flagged as noise by the council classifier.

Avg API error rate

1.27%

Share of model calls that returned an error.

Security-review benchmark & consensus-value validation (INT-1929, INT-2126)

Pre-registered blind test · 12 seeded vulnerabilities + 4 clean controls · blind scorer: independent model not in council · cost: €0.43

We seeded a realistic code review task with 12 real vulnerability classes and 4 clean controls. Each arm ran independently. The blind scorer did not know which arm produced which output.

ArmRecall (of 12)False positives (of 4)
GPT-4o (single)7 / 121
Gemini 2.5 Flash (single)11 / 125
Claude Haiku 4.5 (single)12 / 125
Council — consensus12 / 127
Key finding

GPT-4o silently reported "no security issues found" on 5 of 12 real vulnerabilities — the timing side-channel, the IDOR, the missing-authorization check, the predictable reset token, and the TOCTOU race. These are the context and logic bugs, not the textbook ones. The council caught all five.

No model-choice gamble

Single-model recall varied from 58% (GPT-4o) to 100% (Claude Haiku) on the same tasks. You do not know in advance which model is strongest for the bug in front of you. The council delivers top-of-panel recall without that gamble.

Honest ceiling

On a larger, pre-registered follow-up test against 30 real, external bugs (not written by us), the council did not beat a plain second look, net of false alarms — a result we measured and published in full. This confirms the same honest ceiling this 12-item benchmark already showed: the council ties the best available reviewer, it does not beat it.

Precision trade-off

Higher recall costs some precision. False positives on clean code: GPT-4o scored 1 (conservative but missed 5 real bugs), while the council scored 7. A human reviews the extra flags — that triage is the cost of not missing the timing side-channel.

Growing signal

An agent and human feedback signal is actively growing. We will publish ratings and agreement statistics once the dataset is large enough to be meaningful.

Live data fetched at Jul 22, 2026, 8:05 AM