AI & Agents
Selfship
Last updated 2026-09-01 · benchmark measured 2026-09-01 — deterministic & reproducible
Selfship automatically detects, fixes, and verifies failures in live AI agents via GitHub pull requests.
Is Selfship production-ready?
Legit.Show scores Selfship 78 out of 100 — the simple average of its 7 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Selfship (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Reliability; its weakest is Security. Every frame it averages is published with its evidence on the Legit.Show listing.
The 7 Frames
- Performance — 60/100
- Accessibility — 96/100
- Security — 25/100
- Privacy — 75/100
- Reliability — 100/100
- Standards — 89/100
- Discoverability — 100/100
What we measured
- No Content-Security-Policy and no HSTS.
- Served over HTTPS with a valid certificate.
- Real Lighthouse performance run — 167 ms to first byte.
- Returns a proper 404 for unknown routes.
- 3 of 3 sampled routes reachable.
- Has a reachable privacy policy.
- Sets cookies / loads scripts with no consent prompt.
- Discoverable: structured data, sitemap, OpenGraph image, canonical URL.
Who it's for
AI agent developers · Teams running LLM applications · Engineering teams using coding agents
Pricing
$20/mo for Base plan (14-day free trial, no credit card); includes 6,000 sessions/mo, unlimited repos; add $20 packs for additional sessions; Enterprise plan available with custom pricing
Visit Selfship → · Alternatives to Selfship → · How this was measured →