AI Forecast Ledger
↩ Return to the record

Case file No. 01 · AI Forecast Ledger

On trackcapabilities

A superhuman coder exists by end of 2027

On track · due Dec 2027 On track

3 receiptsOpen verificationClose

On track

SWE-bench Verified %79.20 % resolved on SWE-bench Verified (top agent, bash-only default view) (as of 2026-09-10)

Verified blind: GLM-5.2 + Grok — agree

How it's graded

Met when a model autonomously completes a non-trivial PR end-to-end at senior-eng level

SWE-bench Verified %2 measured points
Measured trajectory for Superhuman coder2 receipt-backed measurements of SWE-bench Verified %, from 79.20 in 2026-08-15 to 79.20 in 2026-09-10. Every value is listed in the table below the figure.78.279.280.2Due Dec 2027Aug 2026Sep 202679.20 % resolved on SWE-bench Verified (top agent, bash-only default view) as of 2026-08-15 — open the source79.20 % resolved on SWE-bench Verified (top agent, bash-only default view) as of 2026-09-10 — open the source79.20 % resolved on SWE-bench Verified (top agent, bash-only default view)
Each point is a dated receipt. Open a point for its archived source.
Table view
As ofMeasured (% resolved on SWE-bench Verified (top agent, bash-only default view))Receipt
2026-08-1579.20source
2026-09-1079.20source

Trajectory unverified indicators — not graded receipts

  • Jun 2025Mid-2025: stumbling agents — first usable AI coding agents appear hitCoding agents emerged in 2025 and now autonomously resolve real GitHub issues on SWE-bench Verified.assessed Jun 2026indication↗
  • Jan 2026Early 2026: coding automation accelerates on paceTop agents now exceed 85% on SWE-bench Verified, up from ~70% a year earlier — fast progress, still short of autonomous senior-eng PRs.assessed Jun 2026indication↗
  • Dec 2027End 2027: a superhuman coder exists (the target) pendingThe claim's target milestone — not yet due.indication↗

Receipts · 3

Ledger history
  • 2026-06-20: seeded (#1)
  • 2026-06-24: replaced placeholder evidence with the SWE-bench Verified leaderboard; verdict held on-track
  • 2026-06-25: demoted pending bulletproof re-grade
  • 2026-06-26: added trajectory checkpoints (AI-2027 milestones, on track) + refreshed SWE-bench measurement to >85%
  • 2026-07-15: Target date is 2027-12 (future). SWE-bench Verified shows top agents above 76% (swebench.com, measurement, SWE-bench), while METR (measurement, different org) confirms AI agents still cannot…
  • 2026-08-15: Verdict on-track remains correct (target 2027-12 is in the future), but a fresh read of the SWE-bench Verified leaderboard (fetched 2026-08-15) shows the top bash-only entry at 79.20% resolved…
  • 2026-09-10: Verdict on-track remains correct (target 2027-12 in the future). A fresh read of the SWE-bench Verified leaderboard (fetched 2026-09-10) confirms the top bash-only entry at 79.20% (Claude 4.5 Opus…

Every verdict on the ledger is graded against dated, archived third-party evidence and blind-verified by two independent models.