Developer Tools
gpt-engineer
Last updated 2026-06-08 · benchmark measured 2026-06-10 — deterministic & reproducible
CLI platform to experiment with codegen. Precursor to: https://lovable.dev
Overall production-readiness score — reserved
Legit.Show benchmarked gpt-engineer across all seven frames. The single overall score is shown to the verified maker; the frame-by-frame breakdown is public below.
Is gpt-engineer production-ready?
Legit.Show ran its deterministic 7-Frame production-readiness benchmark on gpt-engineer (github assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Standards; its weakest is Maintenance. The per-frame breakdown is public on the Legit.Show listing; the single overall production-readiness score is reserved for the verified maker.
The 7 Frames
- Security — 60/100
- Standards — 80/100
- Discoverability — 65/100
- Maintenance — 12/100
What we measured
- No Content-Security-Policy and no HSTS.
- No privacy policy found.
- Sets cookies / loads scripts with no consent prompt.