Legit.Show is a directory of launched web apps, SaaS, AI tools, MCP servers and developer tools — each with an objective 7-Frame production-readiness benchmark, measured deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Methodology · About

Privacy · Terms · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

7-Frame trust gap · 2026 edition

The State of AI-Built Software · 2026

We ran Legit.Show’s 7-Frame production-readiness benchmark across 190 open-source AI, MCP and developer tools — straight from their repositories. AI coding ships a flawless demo; this is what quietly never makes it to production.

96%
ship with no error tracking
across 190 open-source AI & developer tools we measured · according to Legit.Show · 2026

The findings

How many controls are missing

By category

Why these seven

The demo always works — that’s what AI coding is *great* at. The gap is everything a demo never forces you to add: monitoring for when it breaks, limits for when it’s abused, access rules for when there’s more than one user. 1 of 190 (1%) had none of these gaps. The rest are one incident away from finding out.

The basics most get right

It isn’t carelessness with the obvious stuff: 0% shipped a hard-coded secret key in client code, and only 5% committed a `.env`. The misses are the *invisible* controls a human senior adds by reflex and a model rarely does.

What this is not

A health check, not a verdict. A missing rate limit correlates with “shipped fast, hardened never” — it doesn’t prove the product is bad. Every number is a count over a stated sample, measured from the public repository, fully reproducible.

Sample composition

Not a random sample — this is what we measured. The mix below is the caveat; judge it for yourself.

What we measured (190)

The full list, so anyone can spot-check. Every item links to its public benchmark.

How this was measured →