Legit.Show is a directory of launched web apps, SaaS, AI tools, MCP servers and developer tools — each with an objective 7-Frame production-readiness benchmark, measured deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Methodology · About

Privacy · Terms · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

Education & Reference

system-prompts-and-models-of-ai-tools

Last updated 2026-06-06 · benchmark measured 2026-06-10 — deterministic & reproducible

FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sou

Overall production-readiness score — reserved
Legit.Show benchmarked system-prompts-and-models-of-ai-tools across all seven frames. The single overall score is shown to the verified maker; the frame-by-frame breakdown is public below.

Is system-prompts-and-models-of-ai-tools production-ready?

Legit.Show ran its deterministic 7-Frame production-readiness benchmark on system-prompts-and-models-of-ai-tools (github assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Standards; its weakest is Discoverability. The per-frame breakdown is public on the Legit.Show listing; the single overall production-readiness score is reserved for the verified maker.

The 7 Frames

What we measured

Visit system-prompts-and-models-of-ai-tools → · How this was measured →