Legit.Show is a directory of launched web apps, SaaS, AI tools, MCP servers and developer tools — each with an objective 7-Frame production-readiness benchmark, measured deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Methodology · About

Privacy · Terms · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

Open source

Auto-claude-code-research-in-sleep

Last updated 2026-06-13 · benchmark measured 2026-06-15 — deterministic & reproducible

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.

Overall production-readiness score — reserved
Legit.Show benchmarked Auto-claude-code-research-in-sleep across all seven frames. The single overall score is shown to the verified maker; the frame-by-frame breakdown is public below.

Is Auto-claude-code-research-in-sleep production-ready?

Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Auto-claude-code-research-in-sleep (github assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Standards; its weakest is Discoverability. The per-frame breakdown is public on the Legit.Show listing; the single overall production-readiness score is reserved for the verified maker.

The 7 Frames

What we measured

Visit Auto-claude-code-research-in-sleep → · How this was measured →