A German consortium behind the open AI model Soofi S has corrected an error in its evaluation process. The mistake occurred when test questions from the science benchmark GPQA were inadvertently included in the model's training data. The error was discovered by the community after examining publicly available data. The consortium removed the benchmark from its evaluation and recalculated the results. Soofi S remains a top-performing model in both English and German, according to the recalculated results. This correction is crucial for maintaining the integrity of AI model evaluations and ensuring accurate comparisons between different models.