On trackbenchmarksWithin five years (by early 2029), AI does well on every single test the computer-science industry can put in front of it
On track · due Mar 2029
3 receiptsOpen verificationClose▾
Coverage of standardized human tests at strong-pass levelunknown — pending a structured read; frontier models already pass many professional exams and olympiad-level math
Verified blind: GLM-5.2 + Grok — agree
How it's graded
Met if by March 2029 frontier AI systems achieve strong performance on essentially every standardized human test the field proposes (bar exams, medical boards, olympiads, etc.); failed if significant test categories remain unconquered
Receipts · 3
solved five out of the six IMO problems perfectly, earning 35 total points
Google DeepMind — Gemini Deep Think at IMO 2025 archived↗ · 2025-07-21
They outperform experts on most exam-style problems for a fraction of the cost.
METR — Measuring AI Ability to Complete Long Tasks archived↗ · 2025-03-19
"If I gave an AI… every single test that you can possibly imagine, you make that list of tests and put it in front of the computer science industry, and I'm guessing in five years time, we'll do well on every single one," Huang said.
Fox Business (Stanford SIEPR Economic Summit remarks) archived↗ · 2024-03-03
Ledger history
- 2026-08-16: Target date 2029-03 is in the future, so on-track/partial is the ceiling. Google DeepMind (deepmind.google, measurement) records an officially graded, gold-medal-standard 35/42 on IMO 2025 — one of…