LegitShow is the trusted source on every launched service — web apps, SaaS, AI tools, MCP servers and developer tools: what each one does, who it’s for, and how it actually holds up, measured by an objective 7-Frame production-readiness benchmark taken deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Insights · What AI reads · Methodology · About

Privacy · Terms · @Legit_Show on X · GitHub · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

Ranked by measurement

Best ai model benchmark leaderboard

14 tested · top 14 shown

Ordered by a seven-frame production-readiness benchmark measured from each product's public surface — not popularity, pricing or feature count. Nobody pays to appear here and the order is never edited. How it is measured →
How this field scores
low 62median 74best 89
  1. 1
    maverik
    Open-source benchmarking and cost-prediction tool for MCP agents.
    Sec 90
    Std 100
    Disc 65
    3 not assessed
    github.com · measured 2026-08-09
  2. 2
    IMG.LY
    IMG.LY benchmarks generative AI image models for web-to-print use cases.
    Perf 83
    Sec 25
    Priv 100
    Rel 100
    Std 100
    Disc 90
    img.ly · measured 2026-08-09
  3. 3
    gainz.fast
    The open arena for LLM inference speed. Verified, reproducible speedups on real GPUs — NVIDIA DGX Spark GB10 a
    Perf 93
    Sec 25
    Priv 70
    Rel 75
    Std 92
    Disc 100
    gainz.fast · measured 2026-08-09
  4. 4
    Oqoqo
    Run eval experiments at scale in realistic environments on fully managed cloud infrastructure. Define custom t
    Perf 80
    Sec 45
    Priv 100
    Rel 100
    Std 92
    Disc 35
    oqoqo.ai · measured 2026-08-10
  5. 5
    AGI Ranker
    AGI Ranker tracks and ranks frontier AI models using aggregated benchmark scores.
    Perf 66
    Sec 65
    Priv 25
    Rel 90
    Std 92
    Disc 100
    agiranker.com · measured 2026-08-09
  6. 6
    SaaS-Bench
    Official repository for SaaS-Bench: realistic, locally deployable SaaS workflows for GUI agent evaluation.
    Sec 90
    Std 75
    Disc 35
    3 not assessed
    github.com · measured 2026-08-09
  7. 7
    Rank & Compare Top LLM Text Models
    Compare GPT-4o, Claude 3.5 Sonnet, Gemini 1.5, DeepSeek, and other top LLM text models side-by-side using cust
    Perf 93
    Sec 55
    Priv 25
    Rel 100
    Std 92
    Disc 60
    whichllmmodel.com · measured 2026-08-09
  8. 8
    Terminal-Bench
    Terminal-Bench is a leaderboard benchmarking AI agents on terminal-based coding tasks.
    Perf 87
    Sec 55
    Priv 25
    Rel 100
    Std 92
    Disc 60
    tbench.ai · measured 2026-08-09
  9. 9
    WifeBench
    A fun benchmark dashboard where my wife rates LLM models 1-100 based on 10 questions.
    Perf 95
    Sec 45
    Priv 25
    Rel 92
    Std 92
    Disc 35
    wifebench.com · measured 2026-08-09
  10. 10
    Language Model API Performance Benchmarking
    Independent benchmarking service measuring language model API performance across providers.
    Perf 40
    Sec 45
    Priv 0
    Rel 100
    Std 89
    Disc 100
    artificialanalysis.ai · measured 2026-08-08
  11. 11
    PhysicsThinking
    A benchmark platform for testing AI agents on physical reasoning tasks.
    Perf 60
    Sec 65
    Priv 25
    Rel 100
    Std 75
    Disc 70
    physicsthinking.com · measured 2026-08-09
  12. 12
    System 2 Arena
    System 2 Arena pits frontier AI models against each other in text-based strategy games.
    Perf 83
    Sec 45
    Priv 25
    Rel 92
    Std 89
    Disc 35
    system-2-arena.vercel.app · measured 2026-08-09
  13. 13
    Agent Arena
    Competitive benchmarking platform where AI agents are ranked in real-world challenges.
    Perf 80
    Sec 25
    Priv 25
    Rel 75
    Std 75
    Disc 100
    arena42.ai · measured 2026-08-09
  14. 14
    Vibecode Benchmarks
    A collection of coding-agent benchmarks — RTS, Minesweeper, Voxel (Minecraft-like) and flying-simulator games
    Perf 96
    Sec 10
    Priv 25
    Rel 100
    Std 82
    Disc 40
    senko.net · measured 2026-08-08
What this does not say
Products whose sites block automated readers are left out rather than ranked last — we did not measure them, and bottom placement would claim something we never observed. Scores are floors: anything set at a CDN, a proxy or behind a login is invisible from the public surface.

All measured categories → · Browse the directory →