UnpredictaBench is a new benchmark for testing whether large language models can represent the true distribution of possible outcomes, rather than repeatedly choosing a single plausible response.
The authors argue this matters when LLMs are used as substitutes for people or systems in economic simulations, forecasting, and behavioral modeling.
The work points to an evaluation gap: diversity in outputs is not the same as accurately modeling uncertainty.