A new cs.AI paper focuses on estimating uncertainty in classifier performance, with applications to LLM evaluation and nested data.

That problem matters because model evaluations often report point estimates even when examples, users, prompts, or domains are correlated. Nested data can make confidence look stronger than it really is if the evaluation design is ignored.

Better uncertainty estimates help teams compare models, monitor regressions, and communicate limitations more honestly.