The German consortium behind the open 30-billion-parameter Soofi S model has revised its evaluation after benchmark contamination was found in the training data. According to The Decoder, questions from the GPQA science benchmark accidentally appeared in the model’s training set.

The community identified the problem by examining the publicly available data. The team then removed GPQA from its evaluation and recalculated results, a step that matters because contaminated benchmarks can make a model look stronger than it really is.

The case is a reminder that open models need open scrutiny, not just open weights. Public datasets and technical reports allow outside researchers to catch problems that may be invisible in a polished launch announcement.

Soofi S may still be useful, especially for English and German tasks, but the revised results are the relevant ones for comparison. For developers choosing a model, clean evaluation matters as much as headline benchmark rankings.