Highest score on Humanity’s Last Exam in 2026?: how this market works
What you need to know
This market asks how well the best AI model in the world will score on a notoriously difficult test called Humanity's Last Exam by the end of 2026. The test is a collection of extremely hard questions across science, math, and other fields, designed specifically to challenge frontier AI. Three separate yes/no questions are running here: will any AI hit 60% accuracy, 65%, or 70%? A 'Yes' on the 70% question, for example, means some AI correctly answered at least 7 out of every 10 of these very hard questions. Each threshold resolves Yes if any AI model, from any company, scores at least that percentage on the official Humanity's Last Exam leaderboard at agi.safe.ai before midnight Eastern Time on December 31, 2026. The score that counts is specifically the one labeled 'HLE Accuracy' on that leaderboard. Only raw accuracy matters here, not other measures like how confident the model was in its answers. If the leaderboard goes permanently offline and no official alternative exists, all three markets resolve No. The two recent news items provided are not related to AI benchmarks or Humanity's Last Exam, so there is no relevant recent news to point to here. What would matter to watch for: new model releases from major AI labs, any published benchmark runs on HLE, and updates to the official leaderboard at agi.safe.ai. AI capability has been improving fast, but nobody knows exactly how fast. The 60% threshold looks more reachable to the market, priced around 46%, while 70% looks harder at 19%. The core tension is real: HLE was designed to resist AI, but models have surprised experts before. There is also uncertainty about how many labs will even attempt and publish HLE runs before year-end. The gap between the three thresholds shows the market sees a meaningful difference in difficulty between each level.
The odds right now
- 60%+46%
- 65%+31%
- 70%+18%
- 75%+10%
- 80%+8%
- 90%+4%
Price history
60%+
How this resolves
Resolves December 31, 2026
This market will resolve to "Yes" if any model achieves at least the specified accuracy on Humanity’s Last Exam by December 31, 2026, 11:59 PM ET. Otherwise, this market will resolve to "No". Read the full resolution rules on the live market page.
Related
Other outcomes in this market
- 60%+46%
- 65%+31%
- 70%+18%
- 75%+10%
- 80%+8%
- 90%+4%
More markets like this
- Highest Google Gemini score on Humanity’s Last Exam in 2026?
- Highest OpenAI score on Humanity’s Last Exam in 2026?
- Next Claude Opus: Humanity’s Last Exam Debut?
- Highest Grok score on Humanity’s Last Exam in 2026?
- Highest Claude score on Humanity’s Last Exam in 2026?
- Highest Kimi score on Humanity’s Last Exam in 2026?
More market guides
- Highest Google Gemini score on Humanity’s Last Exam in 2026?: how this market works
- Highest OpenAI score on Humanity’s Last Exam in 2026?: how this market works
- Next Claude Opus: Humanity’s Last Exam Debut?: how this market works
- Will any AI model reach … Math Arena Score by December 31?: how this market works
Same markets. A fraction of the fee.
These apps all route to the same exchange order book. The difference is what each one adds on top of the exchange's own fee.
Published rates, checked July 2026. MetaMask charges a flat 4 percent per prediction trade. Jupiter adds a fee equal to the exchange's own taker fee at fill time, roughly 2 to 4 percent at typical odds. Where a market carries an exchange settlement fee, it applies everywhere, whichever app you use. Paridesk adds nothing on maker orders.
Trade this market on Paridesk: non-custodial, 0.5% fee.
View & trade →