Ranked by measurement
Best ai model benchmark leaderboard
14 tested · top 14 shown
Ordered by a seven-frame production-readiness benchmark measured from each product's public surface — not popularity, pricing or feature count. Nobody pays to appear here and the order is never edited. How it is measured →
How this field scores
low 62median 74best 89
- 1
maverikOpen-source benchmarking and cost-prediction tool for MCP agents.Sec 90Std 100Disc 653 not assessedgithub.com · measured 2026-08-09 - 2
IMG.LYIMG.LY benchmarks generative AI image models for web-to-print use cases.Perf 83Sec 25Priv 100Rel 100Std 100Disc 90img.ly · measured 2026-08-09 - 3gainz.fastThe open arena for LLM inference speed. Verified, reproducible speedups on real GPUs — NVIDIA DGX Spark GB10 aPerf 93Sec 25Priv 70Rel 75Std 92Disc 100gainz.fast · measured 2026-08-09
- 4
OqoqoRun eval experiments at scale in realistic environments on fully managed cloud infrastructure. Define custom tPerf 80Sec 45Priv 100Rel 100Std 92Disc 35oqoqo.ai · measured 2026-08-10 - 5
AGI RankerAGI Ranker tracks and ranks frontier AI models using aggregated benchmark scores.Perf 66Sec 65Priv 25Rel 90Std 92Disc 100agiranker.com · measured 2026-08-09 - 6SaaS-BenchOfficial repository for SaaS-Bench: realistic, locally deployable SaaS workflows for GUI agent evaluation.Sec 90Std 75Disc 353 not assessedgithub.com · measured 2026-08-09
- 7Rank & Compare Top LLM Text ModelsCompare GPT-4o, Claude 3.5 Sonnet, Gemini 1.5, DeepSeek, and other top LLM text models side-by-side using custPerf 93Sec 55Priv 25Rel 100Std 92Disc 60whichllmmodel.com · measured 2026-08-09
- 8Terminal-BenchTerminal-Bench is a leaderboard benchmarking AI agents on terminal-based coding tasks.Perf 87Sec 55Priv 25Rel 100Std 92Disc 60tbench.ai · measured 2026-08-09
- 9
WifeBenchA fun benchmark dashboard where my wife rates LLM models 1-100 based on 10 questions.Perf 95Sec 45Priv 25Rel 92Std 92Disc 35wifebench.com · measured 2026-08-09 - 10Language Model API Performance BenchmarkingIndependent benchmarking service measuring language model API performance across providers.Perf 40Sec 45Priv 0Rel 100Std 89Disc 100artificialanalysis.ai · measured 2026-08-08
- 11PhysicsThinkingA benchmark platform for testing AI agents on physical reasoning tasks.Perf 60Sec 65Priv 25Rel 100Std 75Disc 70physicsthinking.com · measured 2026-08-09
- 12
System 2 ArenaSystem 2 Arena pits frontier AI models against each other in text-based strategy games.Perf 83Sec 45Priv 25Rel 92Std 89Disc 35system-2-arena.vercel.app · measured 2026-08-09 - 13Agent ArenaCompetitive benchmarking platform where AI agents are ranked in real-world challenges.Perf 80Sec 25Priv 25Rel 75Std 75Disc 100arena42.ai · measured 2026-08-09
- 14
Vibecode BenchmarksA collection of coding-agent benchmarks — RTS, Minesweeper, Voxel (Minecraft-like) and flying-simulator gamesPerf 96Sec 10Priv 25Rel 100Std 82Disc 40senko.net · measured 2026-08-08
What this does not say
Products whose sites block automated readers are left out rather than ranked last — we did not measure them, and bottom placement would claim something we never observed. Scores are floors: anything set at a CDN, a proxy or behind a login is invisible from the public surface.