How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
Two API settings tripled OpenAI's GPT-5.6 scores on the ARC-AGI-3 benchmark
OpenAI revealed that enabling two specific API settings—retaining reasoning and enabling compaction—tripled GPT-5.6's performance on the ARC-AGI-3 benchmark. This is significant because ARC-AGI-3 is a leading test of general reasoning and fluid intelligence, making this a notable efficiency and capability leap. The finding suggests meaningful headroom in existing models that practitioners may be leaving on the table.