AI & Agents
Referee.chat
Last updated 2026-08-15 · benchmark measured 2026-08-15 — deterministic & reproducible
Referee.chat runs multi-model AI panels to argue a goal until a referee rules it done.
Is Referee.chat production-ready?
Legit.Show scores Referee.chat 93 out of 100 — the simple average of its 7 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Referee.chat (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Privacy; its weakest is Security. Every frame it averages is published with its evidence on the Legit.Show listing.
The 7 Frames
- Performance — 99/100
- Accessibility — 96/100
- Security — 60/100
- Privacy — 100/100
- Reliability — 100/100
- Standards — 97/100
- Discoverability — 100/100
What we measured
- Security headers present: X-Frame-Options, X-Content-Type-Options, Referrer-Policy.
- No Content-Security-Policy and no HSTS.
- Served over HTTPS with a valid certificate.
- Real Lighthouse performance run — 638 ms to first byte.
- Returns a proper 404 for unknown routes.
- 3 of 3 sampled routes reachable.
- Has a reachable privacy policy.
- Sets cookies / loads scripts with no consent prompt.
Who it's for
Researchers · Decision-makers · Mathematicians · Content creators · Professionals seeking evidence-based answers
Pricing
Free tier with $1 starting credit; paid credit from $10 top-up (never expires); charged per-token usage by model; 5% infrastructure fee on BYOK after first $5; unused credit refundable
Visit Referee.chat → · Alternatives to Referee.chat → · How this was measured →