Developer Tools
Trunchbull
Last updated 2026-08-15 · benchmark measured 2026-08-15 — deterministic & reproducible
A platform for benchmarking AI models against tools deployed from GitHub repositories.
Is Trunchbull production-ready?
Legit.Show scores Trunchbull 86 out of 100 — the simple average of its 7 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Trunchbull (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Security; its weakest is Discoverability. Every frame it averages is published with its evidence on the Legit.Show listing.
The 7 Frames
- Performance — 95/100
- Accessibility — 96/100
- Security — 100/100
- Privacy — 100/100
- Reliability — 90/100
- Standards — 85/100
- Discoverability — 35/100
What we measured
- Security headers present: CSP, HSTS, X-Frame-Options, X-Content-Type-Options, Referrer-Policy.
- Served over HTTPS with a valid certificate.
- Real Lighthouse performance run — 710 ms to first byte.
- Returns a proper 404 for unknown routes.
- 2 of 3 sampled routes reachable.
- Has a reachable privacy policy.
- Sets cookies / loads scripts with no consent prompt.
Who it's for
ML engineers · Researchers · Model evaluators · AI developers
Pricing
Free: $0 (1,000 tool calls/month, up to 5 benchmarks); Open Beta: $59.99/month (250,000 tool calls/month, sandbox hours included)
Visit Trunchbull → · Alternatives to Trunchbull → · How this was measured →