2507.04562

Total: 1

#1 Evaluating LLMs on Real-World Forecasting Against Human Superforecasters [PDF] [Copy] [Kimi³] [REL]

Author: Janna Lu

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their ability to forecast future events remains understudied. A year ago, large language models struggle to come close to the accuracy of a human crowd. I evaluate state-of-the-art LLMs on 464 forecasting questions from Metaculus, comparing their performance against human superforecasters. Frontier models achieve Brier scores that ostensibly surpass the human crowd but still significantly underperform a group of superforecasters.

Subjects: Machine Learning , Artificial Intelligence , Computation and Language

Publish: 2025-07-06 22:26:59 UTC

2507.04562

#1 Evaluating LLMs on Real-World Forecasting Against Human Superforecasters [PDF] [Copy] [Kimi3] [REL]

#1 Evaluating LLMs on Real-World Forecasting Against Human Superforecasters [PDF] [Copy] [Kimi³] [REL]