Highest Google Gemini score on Humanity’s Last Exam in 2026?: how this market works
What you need to know
This market asks whether Google's Gemini AI will reach a specific accuracy score on a famously difficult test called Humanity's Last Exam (HLE) before the end of 2026. HLE is a benchmark made up of extremely hard questions across science, math, and other fields, designed specifically to challenge the most advanced AI systems. Three separate score thresholds are tracked: clearing 50% correct, 55% correct, or 60% correct. Right now, getting above half the questions right on this test would already be a significant achievement for any AI. Each threshold resolves Yes if any Google Gemini model, any version, hits that accuracy level on the official HLE leaderboard at agi.safe.ai by December 31, 2026. The score that counts is specifically labeled 'HLE Accuracy', not other related metrics like Calibration Error. One important edge case: if the leaderboard website goes permanently offline and no official alternative source can confirm the score, the market resolves No by default, even if Gemini actually hit the threshold somewhere. None of the recent headlines provided are relevant to this market. They cover unrelated topics like poetry, TikTok trends, and international news. There is no recent news here about Gemini's performance on HLE or any new AI benchmark results. The kind of development worth watching for would be Google announcing a new Gemini model release or an updated HLE leaderboard showing a Gemini score. The three thresholds carry very different levels of confidence: the market prices the 50% threshold at 90%, meaning participants see it as nearly certain, while 60% sits at only 22%, suggesting real doubt. The core difficulty is that AI progress is fast but uneven, and nobody knows exactly when or how much Google will improve Gemini on this specific benchmark. HLE was designed to resist rapid AI progress. The main open question for the 55% and 60% thresholds is simply whether Google's model improvements arrive in time and are large enough.
The odds right now
- 50%+80%
- 55%+41%
- 60%+22%
- 65%+14%
- 70%+10%
Price history
50%+
How this resolves
Resolves December 31, 2026
This market will resolve to "Yes" if any Google Gemini model achieves at least the specified accuracy on Humanity’s Last Exam by December 31, 2026, 11:59 PM ET. Otherwise, this market will resolve to "No". Read the full resolution rules on the live market page.
Related
Other outcomes in this market
- 50%+80%
- 55%+41%
- 60%+22%
- 65%+14%
- 70%+10%
More markets like this
More market guides
Same markets. A fraction of the fee.
These apps all route to the same exchange order book. The difference is what each one adds on top of the exchange's own fee.
Published rates, checked July 2026. MetaMask charges a flat 4 percent per prediction trade. Jupiter adds a fee equal to the exchange's own taker fee at fill time, roughly 2 to 4 percent at typical odds. Where a market carries an exchange settlement fee, it applies everywhere, whichever app you use. Paridesk adds nothing on maker orders.
Trade this market on Paridesk: non-custodial, 0.5% fee.
View & trade →