AI Forecast Ledger

Every AI forecast, graded against reality.

We log the loudest public predictions about AI, then grade each one against dated, linkable evidence. No vibes: only receipts.

The scorecard5 of 12 forecasts settled
1HitIMO gold medal
4MissedRadiologists obsolete · AI writes 90% of code · AI smarter than any human by end of 2025 · AI agents join the workforce in 2025
1 of 5 settled forecasts came true (20%)

7 forecasts still open · 1 on track · 6 pending — not counted until settled

The record

every claim by target date · stamp = current verdict · click a row for the case file
    1. Missed

      Radiologists obsolete

      No. 03Geoffrey Hinton

      Deep learning outperforms radiologists within five years, so we should stop training them now

      Missed

      Missed · called by 2021

      US radiologist workforce headcount vs. 2016US radiologist count rose 17.3% (30,723 to 36,024) from 2014 to 2023, amid a persistent shortage

      Verified blind: DeepSeek V4 + Grok — agree

      How it's graded

      Met if AI had displaced human radiologists (falling demand / headcount) by ~2021; failed if the workforce kept growing and AI became a complement

      Receipts · 3
      Ledger history
      • 2026-06-24: seeded
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-07-02: Target date (2021) long past. Neiman HPI (measurement): radiologist headcount grew 17.3% (30,723 to 36,024) from 2014-2023; Nature (2025) corroborates a persistent shortage. Graded wrong per rubric.

      Open case file No. 03 →

    1. Hit

      IMO gold medal

      No. 02Eliezer Yudkowsky (bet vs. Paul Christiano)

      An AI built before the contest earns a gold medal at the International Mathematical Olympiad by 2025

      Hit

      Hit · resolved by 2025

      IMO score vs. the gold cutoff (35/42 in 2025)Google DeepMind's Gemini Deep Think scored 35/42 — gold-medal standard — at the 2025 IMO

      Verified blind: DeepSeek V4 + Grok — agree

      How it's graded

      Met when, at an IMO through 2025, an AI built beforehand scores at or above the gold-medal cutoff under competition conditions

      Receipts · 3
      • An advanced version of Gemini Deep Think solved five out of the six IMO problems perfectly, earning 35 total points, and achieving gold-medal level performance.

        Google DeepMind archived↗  ·  2025-07-21

      • In July 2025, we reached gold medal-level performance on the International Mathematical Olympiad with a general-purpose reasoning model (35/42 points).

        OpenAI archived↗  ·  2026-02-20

      • So I think we have Paul at <8%, Eliezer at >16% for AI made before the IMO is able to get a gold (under time controls etc. of grand challenge) in one of 2022-2025.

        LessWrong — IMO challenge bet with Eliezer (Paul Christiano) archived↗context · not model-verified  ·  2022-02-25

      Ledger history
      • 2026-06-24: seeded
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-06-25: The claim targets 2025. Google DeepMind's official blog (source 0) confirms an advanced Gemini Deep Think scored 35/42 at IMO 2025, meeting the gold-medal cutoff. OpenAI's First Proof blog (source 1)
      • 2026-06-26: added the Yudkowsky–Christiano bet-odds source (LessWrong, >16% vs <8%) to the receipt; context only — claim, grade, and the 2-model verification binding are unchanged

      Open case file No. 02 →

    2. Missed

      AI smarter than any human by end of 2025

      No. 07Elon MuskDec

      AI smarter than any one human probably around the end of 2025

      Missed

      Missed · called by Dec 2025

      Broad cognitive capability vs. the single most capable humanunknown — pending a formal grade; frontier models lead on many benchmarks but no accepted demonstration of exceeding the smartest human across domains

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if, by end of 2025, a deployed AI system demonstrably exceeds the most capable individual human across the breadth of cognitive tasks (not just benchmarks); failed if leading systems remain below top human experts in substantial domains

      Receipts · 2
      Ledger history
      • 2026-07-05: Target date 2025-12 is past. The claim required AI to exceed the most capable human across the breadth of cognitive tasks by end of 2025. Two recognized-domain measurements show this did not occur…

      Open case file No. 07 →

    3. Missed

      AI agents join the workforce in 2025

      No. 08Sam Altman (OpenAI CEO)Dec

      In 2025, the first AI agents join the workforce and materially change the output of companies

      Missed

      Missed · called by Dec 2025

      Documented production deployments of autonomous agents with material output impactunknown — pending a formal grade; coding agents resolve real GitHub issues autonomously, broader workforce-agent evidence uncollected

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met if by end of 2025 autonomous AI agents are deployed in real companies doing economically material work (documented output changes), beyond copilots and chat assistants; failed if agent deployments remain pilots without material output impact

      Receipts · 2
      Ledger history
      • 2026-07-15: Target date 2025-12 is past. Two recognized-domain measurements confirm AI agents did not materially change company output: METR (metr.org, measurement) states agents cannot carry out substantive…

      Open case file No. 08 →

    1. Missed

      AI writes 90% of code

      No. 04Dario Amodei (Anthropic CEO)Mar

      Within three to six months AI will be writing 90% of code, and within a year essentially all of it

      Missed

      Missed · called by Mar 2026

      Share of production code generated by AI vs. written by humansContested: Amodei cites high internal use at Anthropic, but there is no verified industry-wide 90% figure and independent analyses dispute it

      Verified blind: DeepSeek V4 + Grok — agree

      How it's graded

      Met if, broadly across software, AI is generating ~90% of code by late 2025; failed if human-written code stays the clear majority of production software

      Trajectory unverified indicators — not graded receipts
      • Sep 2025By late 2025: AI writing ~90% of code (the 3–6 month call) missedSix months on, no verified industry-wide 90% figure materialized; independent analyses put AI's share far lower.assessed Jun 2026indication↗
      • Mar 2026By Mar 2026: AI writing essentially all code (the 1-year call) missedA year on, human-written code remains the clear majority of production software; even Anthropic's internal 90% claim is disputed.assessed Jun 2026indication↗
      Receipts · 3
      Ledger history
      • 2026-06-24: seeded
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-06-26: added trajectory checkpoints — both 3–6mo and 1-yr milestones missed
      • 2026-07-03: The claim's target date (2026-03) is in the past. The prediction was that AI would be writing 90% of code by late 2025/early 2026. Source 7 (arXiv, measurement) shows that by December 2024, AI wrote…

      Open case file No. 04 →

    2. Pending

      Agents earn revenue

      No. 06Pundit XDec

      Autonomous agents run measurable revenue by 2026

      Pending

      Pending · due Dec 2026

      attributed revenue %unknown

      How it's graded

      Met when a public company attributes >1% revenue to autonomous agents

      Receipts

      No verified evidence yet.

      Ledger history
      • 2026-06-20: seeded (#1)

      Open case file No. 06 →

  1. Today · Aug 2026
    1. On track

      Superhuman coder

      No. 01AI-2027 reportDec

      A superhuman coder exists by end of 2027

      On track

      On track · due Dec 2027

      SWE-bench Verified %Top coding agents now exceed 85% on SWE-bench Verified (mid-2026), up from ~70% a year earlier — still short of autonomous senior-eng PRs

      Verified blind: GLM-5.2 + Grok — agree

      How it's graded

      Met when a model autonomously completes a non-trivial PR end-to-end at senior-eng level

      Trajectory unverified indicators — not graded receipts
      • Jun 2025Mid-2025: stumbling agents — first usable AI coding agents appear hitCoding agents emerged in 2025 and now autonomously resolve real GitHub issues on SWE-bench Verified.assessed Jun 2026indication↗
      • Jan 2026Early 2026: coding automation accelerates on paceTop agents now exceed 85% on SWE-bench Verified, up from ~70% a year earlier — fast progress, still short of autonomous senior-eng PRs.assessed Jun 2026indication↗
      • Dec 2027End 2027: a superhuman coder exists (the target) pendingThe claim's target milestone — not yet due.indication↗
      Receipts · 2
      Ledger history
      • 2026-06-20: seeded (#1)
      • 2026-06-24: replaced placeholder evidence with the SWE-bench Verified leaderboard; verdict held on-track
      • 2026-06-25: demoted pending bulletproof re-grade
      • 2026-06-26: added trajectory checkpoints (AI-2027 milestones, on track) + refreshed SWE-bench measurement to >85%
      • 2026-07-15: Target date is 2027-12 (future). SWE-bench Verified shows top agents above 76% (swebench.com, measurement, SWE-bench), while METR (measurement, different org) confirms AI agents still cannot…

      Open case file No. 01 →

    1. Pending

      Weakly-general AI

      No. 05Metaculus community forecast

      The first weakly general AI system is publicly announced around 2028

      Pending

      Pending · due 2028

      Public announcement date vs. the live community medianunknown — pending a direct read of the live Metaculus median

      How it's graded

      Met if a system meeting the Metaculus weakly-general-AI resolution criteria is publicly announced near the community median

      Receipts

      No verified evidence yet.

      Ledger history
      • 2026-06-24: seeded; awaiting a verifiable read of the Metaculus median before grading

      Open case file No. 05 →

    1. Pending

      Turing test passed by 2029

      No. 09Ray Kurzweil

      AI reaches human-level intelligence and passes the Turing test by 2029

      Pending

      Pending · due 2029

      Rigorous adversarial Turing-test passunknown — no accepted rigorous adversarial Turing-test pass on record

      How it's graded

      Met if before end of 2029 an AI passes a rigorous adversarial Turing test (expert judges, extended sessions) or an equivalent accepted human-level-intelligence demonstration; failed otherwise

      Receipts · 1

      Open case file No. 09 →

    2. Pending

      AI passes every human test by 2029

      No. 10Jensen Huang (Nvidia CEO)Mar

      Within five years (by early 2029), AI does well on every single test the computer-science industry can put in front of it

      Pending

      Pending · due Mar 2029

      Coverage of standardized human tests at strong-pass levelunknown — pending a structured read; frontier models already pass many professional exams and olympiad-level math

      How it's graded

      Met if by March 2029 frontier AI systems achieve strong performance on essentially every standardized human test the field proposes (bar exams, medical boards, olympiads, etc.); failed if significant test categories remain unconquered

      Receipts · 1
      • "If I gave an AI… every single test that you can possibly imagine, you make that list of tests and put it in front of the computer science industry, and I'm guessing in five years time, we'll do well on every single one," Huang said.

        Fox Business (Stanford SIEPR Economic Summit remarks)  ·  2024-03-03

      Open case file No. 10 →

    1. Pending

      LeCun: human-level AI a decade out

      No. 12Yann LeCun (Meta Chief AI Scientist)

      Human-level AI is years away, if not a decade — and will not come from LLMs alone (counterpoint claim)

      Pending

      Pending · due 2034

      Arrival date and architecture of the first accepted human-level systemunknown — long-dated counterpoint, tracking

      How it's graded

      Met if human-level AI does NOT appear before ~2030 (vindicating the skeptic call); failed if a widely accepted human-level system arrives well before the end of the decade or emerges primarily from LLM scaling

      Receipts · 1
      • “It’s going to take years before we can get everything here to work, if not a decade,” said LeCun. “Mark Zuckerberg keeps asking me how long it’s going to take.”

        TechCrunch (Hudson Forum keynote)  ·  2024-10-16

      Open case file No. 12 →

    1. Pending

      Hassabis: AGI in 5–10 years

      No. 11Demis Hassabis (Google DeepMind CEO)

      AGI arrives in the next five to ten years (2030–2035)

      Pending

      Pending · due 2035

      Public demonstration of a generally human-level systemunknown — long-dated claim, tracking

      How it's graded

      Met if a system widely accepted as AGI (general human-level capability across domains) is publicly demonstrated between 2030 and 2035; graded early as wrong only if Hassabis's stated criteria are clearly unmet by end of 2035

      Receipts · 1

      Open case file No. 11 →

How grading works
Hitthe prediction came true on its terms. On tracktrending toward true before its date. At risktrending toward false, or contested. Missedits date passed and it did not come true. Overduedeadline passed, not yet graded. Pendingnot yet gradable, or evidence still unverified.

Open any dossier for its receipts. Logged from public predictions → graded against dated, linkable evidence → every receipt an archived third-party source, drafted by an agent and approved by a human.

Field Notes

Field Notes — commentary, not a graded receipt

How claims are verified

Every resolved claim requires at least 2 independent sources — each with an archived snapshot — before a verdict is assigned. Grading is done blind: two independent models each evaluate the claim without seeing the other's judgement; both must agree. Each verdict names the exact models that graded it. If they disagree the claim is marked contested and excluded from all headline stats until a human reviewer resolves it.

Every claim's text, sources, and grade history are pinned in an append-only ledger; grade integrity is enforced by the validation gate and the per-model verification records.

Unverified and contested claims appear on the board for transparency but are not counted in the "settled forecasts came true" figure. Trajectory checkpoints are dated indications of progress, not graded verdicts — they never affect the "settled forecasts came true" figure.