On trackcapabilitiesA superhuman coder exists by end of 2027
On track · due Dec 2027 On track
2 receiptsOpen verificationClose▾
On track · due Dec 2027
SWE-bench Verified %Top coding agents now exceed 85% on SWE-bench Verified (mid-2026), up from ~70% a year earlier — still short of autonomous senior-eng PRs
Verified blind: GLM-5.2 + Grok — agree
Met when a model autonomously completes a non-trivial PR end-to-end at senior-eng level
- Jun 2025Mid-2025: stumbling agents — first usable AI coding agents appear hitCoding agents emerged in 2025 and now autonomously resolve real GitHub issues on SWE-bench Verified.assessed Jun 2026indication↗
- Jan 2026Early 2026: coding automation accelerates on paceTop agents now exceed 85% on SWE-bench Verified, up from ~70% a year earlier — fast progress, still short of autonomous senior-eng PRs.assessed Jun 2026indication↗
- Dec 2027End 2027: a superhuman coder exists (the target) pendingThe claim's target milestone — not yet due.indication↗
Claude 4.5 Opus (high reasoning) | 76.80
SWE-bench Verified leaderboard archived↗ · 2026-07-15
best AI agents are not currently able to carry out substantive projects by themselves or directly substitute for human labor.
METR — Measuring AI Ability to Complete Long Tasks archived↗ · 2025-03-19
Ledger history
- 2026-06-20: seeded (#1)
- 2026-06-24: replaced placeholder evidence with the SWE-bench Verified leaderboard; verdict held on-track
- 2026-06-25: demoted pending bulletproof re-grade
- 2026-06-26: added trajectory checkpoints (AI-2027 milestones, on track) + refreshed SWE-bench measurement to >85%
- 2026-07-15: Target date is 2027-12 (future). SWE-bench Verified shows top agents above 76% (swebench.com, measurement, SWE-bench), while METR (measurement, different org) confirms AI agents still cannot…