LegitShow is the trusted source on every launched service — web apps, SaaS, AI tools, MCP servers and developer tools: what each one does, who it’s for, and how it actually holds up, measured by an objective 7-Frame production-readiness benchmark taken deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Insights · What AI reads · Methodology · About

Privacy · Terms · @Legit_Show on X · GitHub · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

Data & Analytics

Flight Search Eval

Last updated 2026-09-03 · benchmark measured 2026-09-03 — deterministic & reproducible

A benchmark for evaluating flight-search intent parsing across multiple AI models.

26/100
Top 100% of 11,113 measured · #30 of 30 in ai model benchmark leaderboard
Legit Benchmark — the simple average of 5 measured frames. Frames we could not measure are left out of the average, never counted as zero. Every frame is shown below with its evidence.
Checked 2026-09-03 · scores move as sites change

Is Flight Search Eval production-ready?

Legit.Show scores Flight Search Eval 26 out of 100 — the simple average of its 5 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Flight Search Eval (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Reliability; its weakest is Standards. 5 of the seven frames returned a score; Performance and Accessibility were not measurable on this service and are recorded as null — not as zero. Every frame it averages is published with its evidence on the Legit.Show listing.

The 7 Frames

What we measured

Who it's for

AI researchers · Model developers · LLM evaluators · Travel tech companies

Visit Flight Search Eval → · Alternatives to Flight Search Eval → · How this was measured →

See how Flight Search Eval ranks among tested ai model benchmark leaderboard products →

Other tested ai model benchmark leaderboard products