The market price of reaching a fixed score on selected AI benchmarks has fallen rapidly since 2023, but that does not mean today’s most capable models are becoming equally cheap. Epoch AI estimates an average decline of about 47% per quarter, or roughly 13-fold per year, when comparing the price of a constant performance level.
As an example, Epoch says OpenAI’s o3 reached 75% on the GPQA Diamond science test in early 2025 at an estimated 30 cents per question. Eighteen months later, a GPT-5.6 family model reportedly matched that score for four-hundredths of a cent. Epoch calls its estimate rough because it draws on only five math, science and logic benchmarks.
MIT researchers examining broader pricing data found annual declines of five- to tenfold. After separating cheaper hardware and competition, they estimated algorithmic efficiency improved about threefold per year. The analyses answer different questions, so their headline rates are not directly interchangeable.
New frontier models can still cost more per answer because reasoning systems spend additional compute on difficult tasks. Benchmarks may also reward targeted optimization that does not transfer to practical work. For buyers, price per token or fixed benchmark score is only part of the decision; latency, errors, retries, context limits and output speed can outweigh the cheapest listed rate.