Highest Grok score on Humanity’s Last Exam in 2026?: how this market works
What you need to know
This market asks how well xAI's Grok AI model will score on a very hard academic test called Humanity's Last Exam (HLE) before the end of 2026. HLE is a benchmark made up of extremely difficult questions across science, math, and other fields, designed to challenge the most advanced AI systems. There are three separate versions of this market at different score thresholds: 45%, 50%, and 55% accuracy. A 'Yes' means Grok crosses that threshold; a 'No' means it falls short. Each version resolves Yes if any Grok model from xAI appears on the official HLE leaderboard at agi.safe.ai with an accuracy score at or above its specific threshold (45%, 50%, or 55%) before December 31, 2026 at 11:59 PM Eastern Time. The score must be labeled 'HLE Accuracy' or a clear equivalent. Importantly, only the raw accuracy number counts, not the model's Calibration Error or other related metrics. If the leaderboard goes permanently offline and no official alternative exists, all versions resolve No. None of the provided news headlines relate to Grok, xAI, or Humanity's Last Exam. There is no relevant recent news to draw from here. The kind of developments worth watching for would be: xAI announcing a new Grok model release, any published HLE benchmark results featuring Grok, or updates to the HLE leaderboard showing new entries. AI benchmark progress is genuinely hard to forecast because model improvements can be sudden and large, and xAI's release schedule is not public. The market currently prices the 45% threshold at 53%, meaning it sees the outcome as roughly a coin flip, while 55% is priced at only 14%, reflecting how much harder that bar is. The main unknowns are how many new Grok versions xAI releases before the deadline, how much each improves, and whether HLE scores scale predictably with model size and training.
The odds right now
- 45%+46%
- 50%+32%
- 55%+15%
- 60%+8%
- 65%+5%
Price history
45%+
How this resolves
Resolves December 31, 2026
This market will resolve to "Yes" if any SpaceXAI Grok model achieves at least the specified accuracy on Humanity’s Last Exam by December 31, 2026, 11:59 PM ET. Otherwise, this market will resolve to "No". Read the full resolution rules on the live market page.
Related
Other outcomes in this market
- 45%+46%
- 50%+32%
- 55%+15%
- 60%+8%
- 65%+5%
More markets like this
- Will any AI model reach … Overall Arena Score by September 30?
- Highest Google Gemini score on Humanity’s Last Exam in 2026?
- Will any AI model reach … Coding Arena Score by December 31?
- Next Claude Opus: Humanity’s Last Exam Debut?
- Highest score on Humanity’s Last Exam in 2026?
- Will any AI model reach … Overall Arena Score by December 31?
More market guides
- Will any AI model reach … Overall Arena Score by September 30?: how this market works
- Highest Google Gemini score on Humanity’s Last Exam in 2026?: how this market works
- Will any AI model reach … Coding Arena Score by December 31?: how this market works
- Next Claude Opus: Humanity’s Last Exam Debut?: how this market works
Same markets. A fraction of the fee.
These apps all route to the same exchange order book. The difference is what each one adds on top of the exchange's own fee.
Published rates, checked July 2026. MetaMask charges a flat 4 percent per prediction trade. Jupiter adds a fee equal to the exchange's own taker fee at fill time, roughly 2 to 4 percent at typical odds. Where a market carries an exchange settlement fee, it applies everywhere, whichever app you use. Paridesk adds nothing on maker orders.
Trade this market on Paridesk: non-custodial, 0.5% fee.
View & trade →