AI & Agents
Bench-bench
Last updated 2026-08-18 · benchmark measured 2026-08-18 — deterministic & reproducible
Bench-bench benchmarks AI models by simulating a weightlifting training program.
Is Bench-bench production-ready?
Legit.Show scores Bench-bench 58 out of 100 — the simple average of its 7 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Bench-bench (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Reliability; its weakest is Privacy. Every frame it averages is published with its evidence on the Legit.Show listing.
The 7 Frames
- Performance — 55/100
- Accessibility — 90/100
- Security — 25/100
- Privacy — 25/100
- Reliability — 92/100
- Standards — 67/100
- Discoverability — 50/100
What we measured
- No Content-Security-Policy and no HSTS.
- Served over HTTPS with a valid certificate.
- Real Lighthouse performance run — 105 ms to first byte.
- Returns a proper 404 for unknown routes.
- 0 of 0 sampled routes reachable.
- No privacy policy found.
- Sets cookies / loads scripts with no consent prompt.
- Discoverable: canonical URL.
Who it's for
AI researchers · Model developers · AI enthusiasts · Fitness enthusiasts · Benchmark researchers
Pricing
$115 for v0.2 public run (4 models, 40 simulated years)
Visit Bench-bench → · Alternatives to Bench-bench → · How this was measured →