A new cs.CL paper introduces Know2Guess, a contamination-aware benchmark for measuring the boundary between answerable knowledge and cases where a language model should abstain.
That distinction is important because many evaluations blur several failure modes together: the model may know the answer, have seen the data during training, guess from weak clues, or refuse in a generic way. A benchmark that separates those zones can make reliability claims more meaningful.
The work fits a larger push toward evaluations that test calibrated behavior, not only raw accuracy on static question sets.