AI Forecast Ledger

Every AI forecast, graded against reality.

We log the loudest public predictions about AI, then grade each one against dated, linkable evidence. No vibes: only receipts.

The scorecard

5 of 11 forecasts settled
1HitIMO gold medal
4MissedRadiologists obsolete · AI writes 90% of code · AI smarter than any human by end of 2025 · AI agents join the workforce in 2025
1 of 5 settled forecasts came true (20%)

6 forecasts still open · 5 on track · 1 pending — not counted until settled

today20212023202520272029203120332035Missed — Radiologists obsoleteMissed · Radiologists obsolete · 2021Hit — IMO gold medalHit · IMO gold medal · 2025Missed — AI smarter than any human by end of 2025Missed · AI smarter than any human by end of 2025 · Dec 2025Missed — AI agents join the workforce in 2025Missed · AI agents join the workforce in 2025 · Dec 2025Missed — AI writes 90% of codeMissed · AI writes 90% of code · Mar 2026On track — Superhuman coderOn track · Superhuman coder · Dec 2027Pending — Weakly-general AIPending · Weakly-general AI · 2028On track — AI passes every human test by 2029On track · AI passes every human test by 2029 · Mar 2029On track — Turing test passed by 2029On track · Turing test passed by 2029 · 2029On track — LeCun: human-level AI a decade outOn track · LeCun: human-level AI a decade out · 2034On track — Hassabis: AGI in 5–10 yearsOn track · Hassabis: AGI in 5–10 years · 2035
  • Hit
  • Missed
  • On track
  • Pending

The record

every claim by target date · stamp = current verdict · click a row to open its dossier
    1. Missed

      Radiologists obsolete

      No. 03Geoffrey Hintongraded Sep 2026

      Deep learning outperforms radiologists within five years, so we should stop training them now

      Missed

      US radiologist workforce headcount vs. 201617.3 % growth in US radiologist headcount, 2014-2023 (as of 2023)

      Verified blind: DeepSeek V4 + Grok — agree

      How it's graded

      Met if AI had displaced human radiologists (falling demand / headcount) by ~2021; failed if the workforce kept growing and AI became a complement

      US radiologist workforce headcount vs. 20162 measured points
      Measured trajectory for Radiologists obsolete2 receipt-backed measurements of US radiologist workforce headcount vs. 2016, from 17.3 in 2023 to 25.7 in 2025-02-12. Every value is listed in the table below the figure.1721.526Due 20212023Feb 202517.3 % growth in US radiologist headcount, 2014-2023 as of 2023 — open the source25.7 % growth in US radiologist headcount, 2014-2023 as of 2025-02-12 — open the source25.7 % growth in US radiologist headcount, 2014-2023
      Each point is a dated receipt. Open a point for its archived source.
      Table view
      As ofMeasured (% growth in US radiologist headcount, 2014-2023)Receipt
      202317.3source
      2025-02-1225.7source
      Receipts · 3
      Ledger history
      • 2026-06-24: seeded
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-07-02: Target date (2021) long past. Neiman HPI (measurement): radiologist headcount grew 17.3% (30,723 to 36,024) from 2014-2023; Nature (2025) corroborates a persistent shortage. Graded wrong per rubric.
      • 2026-08-17: Verdict 'wrong' remains correct: the 2021 target is long past and the radiologist workforce grew rather than collapsed. A fresher Neiman HPI release (Feb 2025, neimanhpi.org) refreshes the workforce…
      • 2026-09-14: Verdict 'wrong' remains correct (2021 target long past; the workforce grew rather than collapsed). A fresh read of the Neiman HPI release (fetched 2026-09-14) provides a concrete headcount…

      Open case file No. 03 →

    1. Hit

      IMO gold medal

      No. 02Eliezer Yudkowsky (bet vs. Paul Christiano)graded Jun 2026

      An AI built before the contest earns a gold medal at the International Mathematical Olympiad by 2025

      Hit

      IMO score vs. the gold cutoff (35/42 in 2025)Google DeepMind's Gemini Deep Think scored 35/42 — gold-medal standard — at the 2025 IMO

      Verified blind: DeepSeek V4 + Grok — agree

      How it's graded

      Met when, at an IMO through 2025, an AI built beforehand scores at or above the gold-medal cutoff under competition conditions

      Receipts · 3
      • An advanced version of Gemini Deep Think solved five out of the six IMO problems perfectly, earning 35 total points, and achieving gold-medal level performance.

        Google DeepMind archived↗  ·  2025-07-21

      • In July 2025, we reached gold medal-level performance on the International Mathematical Olympiad with a general-purpose reasoning model (35/42 points).

        OpenAI archived↗  ·  2026-02-20

      • So I think we have Paul at <8%, Eliezer at >16% for AI made before the IMO is able to get a gold (under time controls etc. of grand challenge) in one of 2022-2025.

        LessWrong — IMO challenge bet with Eliezer (Paul Christiano) archived↗context · not model-verified  ·  2022-02-25

      Ledger history
      • 2026-06-24: seeded
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-06-25: The claim targets 2025. Google DeepMind's official blog (source 0) confirms an advanced Gemini Deep Think scored 35/42 at IMO 2025, meeting the gold-medal cutoff. OpenAI's First Proof blog (source 1)
      • 2026-06-26: added the Yudkowsky–Christiano bet-odds source (LessWrong, >16% vs <8%) to the receipt; context only — claim, grade, and the 2-model verification binding are unchanged

      Open case file No. 02 →

    2. Missed

      AI smarter than any human by end of 2025

      No. 06Elon MuskDecgraded Jul 2026

      AI smarter than any one human probably around the end of 2025

      Missed

      Broad cognitive capability vs. the single most capable humanunknown — pending a formal grade; frontier models lead on many benchmarks but no accepted demonstration of exceeding the smartest human across domains

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if, by end of 2025, a deployed AI system demonstrably exceeds the most capable individual human across the breadth of cognitive tasks (not just benchmarks); failed if leading systems remain below top human experts in substantial domains

      Receipts · 3
      Ledger history
      • 2026-07-05: Target date 2025-12 is past. The claim required AI to exceed the most capable human across the breadth of cognitive tasks by end of 2025. Two recognized-domain measurements show this did not occur…

      Open case file No. 06 →

    3. Missed

      AI agents join the workforce in 2025

      No. 07Sam Altman (OpenAI CEO)Decgraded Jul 2026

      In 2025, the first AI agents join the workforce and materially change the output of companies

      Missed

      Documented production deployments of autonomous agents with material output impactunknown — pending a formal grade; coding agents resolve real GitHub issues autonomously, broader workforce-agent evidence uncollected

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if by end of 2025 autonomous AI agents are deployed in real companies doing economically material work (documented output changes), beyond copilots and chat assistants; failed if agent deployments remain pilots without material output impact

      Receipts · 3
      Ledger history
      • 2026-07-15: Target date 2025-12 is past. Two recognized-domain measurements confirm AI agents did not materially change company output: METR (metr.org, measurement) states agents cannot carry out substantive…

      Open case file No. 07 →

    1. Missed

      AI writes 90% of code

      No. 04Dario Amodei (Anthropic CEO)Margraded Sep 2026

      Within three to six months AI will be writing 90% of code, and within a year essentially all of it

      Missed

      Share of production code generated by AI vs. written by humans30.1 % of Python functions (US contributors) written by AI, December 2024 (as of 2024-12)

      Verified blind: DeepSeek V4 + Grok — agree

      How it's graded

      Met if, broadly across software, AI is generating ~90% of code by late 2025; failed if human-written code stays the clear majority of production software

      Share of production code generated by AI vs. written by humans1 measured point

      30.1% of Python functions (US contributors) written by AI, December 2024as of 2024-12

      A single reading is not yet a trajectory. The figure appears once a second measurement lands.

      Table view
      As ofMeasured (% of Python functions (US contributors) written by AI, December 2024)Receipt
      2024-1230.1source
      Trajectory unverified indicators — not graded receipts
      • Sep 2025By late 2025: AI writing ~90% of code (the 3–6 month call) missedSix months on, no verified industry-wide 90% figure materialized; independent analyses put AI's share far lower.assessed Jun 2026indication↗
      • Mar 2026By Mar 2026: AI writing essentially all code (the 1-year call) missedA year on, human-written code remains the clear majority of production software; even Anthropic's internal 90% claim is disputed.assessed Jun 2026indication↗
      Receipts · 3
      Ledger history
      • 2026-06-24: seeded
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-06-26: added trajectory checkpoints — both 3–6mo and 1-yr milestones missed
      • 2026-07-03: The claim's target date (2026-03) is in the past. The prediction was that AI would be writing 90% of code by late 2025/early 2026. Source 7 (arXiv, measurement) shows that by December 2024, AI wrote…
      • 2026-09-14: Verdict 'wrong' remains correct (target 2026-03 passed with AI's share of production code far below 90%). A fresh read of the arXiv study supplies the best quantified independent measurement of…

      Open case file No. 04 →

  1. Today · Sep 2026
    1. On track

      Superhuman coder

      No. 01AI-2027 reportDec1y 3mo on the clock

      A superhuman coder exists by end of 2027

      On track

      SWE-bench Verified %79.20 % resolved on SWE-bench Verified (top agent, bash-only default view) (as of 2026-09-10)

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met when a model autonomously completes a non-trivial PR end-to-end at senior-eng level

      SWE-bench Verified %2 measured points
      Measured trajectory for Superhuman coder2 receipt-backed measurements of SWE-bench Verified %, from 79.20 in 2026-08-15 to 79.20 in 2026-09-10. Every value is listed in the table below the figure.78.279.280.2Due Dec 2027Aug 2026Sep 202679.20 % resolved on SWE-bench Verified (top agent, bash-only default view) as of 2026-08-15 — open the source79.20 % resolved on SWE-bench Verified (top agent, bash-only default view) as of 2026-09-10 — open the source79.20 % resolved on SWE-bench Verified (top agent, bash-only default view)
      Each point is a dated receipt. Open a point for its archived source.
      Table view
      As ofMeasured (% resolved on SWE-bench Verified (top agent, bash-only default view))Receipt
      2026-08-1579.20source
      2026-09-1079.20source
      Trajectory unverified indicators — not graded receipts
      • Jun 2025Mid-2025: stumbling agents — first usable AI coding agents appear hitCoding agents emerged in 2025 and now autonomously resolve real GitHub issues on SWE-bench Verified.assessed Jun 2026indication↗
      • Jan 2026Early 2026: coding automation accelerates on paceTop agents now exceed 85% on SWE-bench Verified, up from ~70% a year earlier — fast progress, still short of autonomous senior-eng PRs.assessed Jun 2026indication↗
      • Dec 2027End 2027: a superhuman coder exists (the target) pendingThe claim's target milestone — not yet due.indication↗
      Receipts · 3
      Ledger history
      • 2026-06-20: seeded (#1)
      • 2026-06-24: replaced placeholder evidence with the SWE-bench Verified leaderboard; verdict held on-track
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-06-26: added trajectory checkpoints (AI-2027 milestones, on track) + refreshed SWE-bench measurement to >85%
      • 2026-07-15: Target date is 2027-12 (future). SWE-bench Verified shows top agents above 76% (swebench.com, measurement, SWE-bench), while METR (measurement, different org) confirms AI agents still cannot…
      • 2026-08-15: Verdict on-track remains correct (target 2027-12 is in the future), but a fresh read of the SWE-bench Verified leaderboard (fetched 2026-08-15) shows the top bash-only entry at 79.20% resolved…
      • 2026-09-10: Verdict on-track remains correct (target 2027-12 in the future). A fresh read of the SWE-bench Verified leaderboard (fetched 2026-09-10) confirms the top bash-only entry at 79.20% (Claude 4.5 Opus…

      Open case file No. 01 →

    1. Pending

      Weakly-general AI

      No. 05Metaculus community forecast2y 3mo on the clock

      The first weakly general AI system is publicly announced around 2028

      Pending

      Public announcement date vs. the live community medianunknown — pending a direct read of the live Metaculus median

      How it's graded

      Met if a system meeting the Metaculus weakly-general-AI resolution criteria is publicly announced near the community median

      Receipts

      No verified evidence yet.

      Ledger history
      • 2026-06-24: seeded; awaiting a verifiable read of the Metaculus median before grading

      Open case file No. 05 →

    1. On track

      Turing test passed by 2029

      No. 08Ray Kurzweil3y 3mo on the clock

      AI reaches human-level intelligence and passes the Turing test by 2029

      On track

      Rigorous adversarial Turing-test passunknown — no accepted rigorous adversarial Turing-test pass on record

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if before end of 2029 an AI passes a rigorous adversarial Turing test (expert judges, extended sessions) or an equivalent accepted human-level-intelligence demonstration; failed otherwise

      Receipts · 3
      Ledger history
      • 2026-08-17: Target date 2029 is in the future, so the verdict is capped at on-track/partial. Two recognized-domain evidence items with different origin_orgs (>=1 measurement) document concrete positive progress…

      Open case file No. 08 →

    2. On track

      AI passes every human test by 2029

      No. 09Jensen Huang (Nvidia CEO)Mar2y 6mo on the clock

      Within five years (by early 2029), AI does well on every single test the computer-science industry can put in front of it

      On track

      Coverage of standardized human tests at strong-pass levelunknown — pending a structured read; frontier models already pass many professional exams and olympiad-level math

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if by March 2029 frontier AI systems achieve strong performance on essentially every standardized human test the field proposes (bar exams, medical boards, olympiads, etc.); failed if significant test categories remain unconquered

      Receipts · 3
      Ledger history
      • 2026-08-16: Target date 2029-03 is in the future, so on-track/partial is the ceiling. Google DeepMind (deepmind.google, measurement) records an officially graded, gold-medal-standard 35/42 on IMO 2025 — one of…

      Open case file No. 09 →

    1. On track

      LeCun: human-level AI a decade out

      No. 11Yann LeCun (Meta Chief AI Scientist)8y 3mo on the clock

      Human-level AI is years away, if not a decade — and will not come from LLMs alone (counterpoint claim)

      On track

      Arrival date and architecture of the first accepted human-level systemunknown — long-dated counterpoint, tracking

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if human-level AI does NOT appear before ~2030 (vindicating the skeptic call); failed if a widely accepted human-level system arrives well before the end of the decade or emerges primarily from LLM scaling

      Receipts · 3
      Ledger history
      • 2026-09-10: Target date 2034 is far in the future, so on-track/partial is the ceiling. LeCun's skeptic call is currently holding: METR (metr.org, measurement, different org) measures that the best agents still…

      Open case file No. 11 →

    1. On track

      Hassabis: AGI in 5–10 years

      No. 10Demis Hassabis (Google DeepMind CEO)9y 3mo on the clock

      AGI arrives in the next five to ten years (2030–2035)

      On track

      Public demonstration of a generally human-level systemunknown — long-dated claim, tracking

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if a system widely accepted as AGI (general human-level capability across domains) is publicly demonstrated between 2030 and 2035; graded early as wrong only if Hassabis's stated criteria are clearly unmet by end of 2035

      Receipts · 3
      Ledger history
      • 2026-09-03: Target date 2035 is well in the future, so on-track/partial is the ceiling. Two recognized-domain evidence items with different origin_orgs (both role:measurement) show concrete capability progress…

      Open case file No. 10 →

How grading works
Hitthe prediction came true on its terms. On tracktrending toward true before its date. At risktrending toward false, or contested. Missedits date passed and it did not come true. Overduedeadline passed, not yet graded. Pendingnot yet gradable, or evidence still unverified.

Open any dossier for its receipts. Logged from public predictions → graded against dated, linkable evidence → every receipt an archived third-party source, drafted by an agent and approved by a human.

Field Notes

Field Notes — commentary, not a graded receipt

How claims are verified

Every resolved claim requires at least 2 independent sources — each with an archived snapshot — before a verdict is assigned. Grading is done blind: two independent models each evaluate the claim without seeing the other's judgement; both must agree. Each verdict names the exact models that graded it. If they disagree the claim is marked contested and excluded from all headline stats until a human reviewer resolves it.

Every claim's text, sources, and grade history are pinned in an append-only ledger; grade integrity is enforced by the validation gate and the per-model verification records.

Unverified and contested claims appear on the board for transparency but are not counted in the "settled forecasts came true" figure. Trajectory checkpoints are dated indications of progress, not graded verdicts — they never affect the "settled forecasts came true" figure.