OpenAI has reported that its GPT-5.6 Sol model outperformed Anthropic's Opus 5 on the ARC-AGI-3 test, achieving a score of 38.3 percent. However, this result was obtained using OpenAI's own custom test environment, which includes features such as retained reasoning and context compaction. In contrast, Opus 5 achieved a score of 30.2 percent without these aids. When tested in the official environment, GPT-5.6 Sol's score dropped to 7.8 percent. The discrepancy highlights the importance of standardized testing conditions in evaluating AI models.