AI Forecast Ledger
↩ Return to the record

Case file No. 09 · AI Forecast Ledger

On trackbenchmarks

Within five years (by early 2029), AI does well on every single test the computer-science industry can put in front of it

On track · due Mar 2029

3 receiptsOpen verificationClose

On track

Coverage of standardized human tests at strong-pass levelunknown — pending a structured read; frontier models already pass many professional exams and olympiad-level math

Verified blind: GLM-5.2 + Grok — agree

How it's graded

Met if by March 2029 frontier AI systems achieve strong performance on essentially every standardized human test the field proposes (bar exams, medical boards, olympiads, etc.); failed if significant test categories remain unconquered

Receipts · 3

Ledger history
  • 2026-08-16: Target date 2029-03 is in the future, so on-track/partial is the ceiling. Google DeepMind (deepmind.google, measurement) records an officially graded, gold-medal-standard 35/42 on IMO 2025 — one of…

Every verdict on the ledger is graded against dated, archived third-party evidence and blind-verified by two independent models.