METR found OpenAI GPT-5.6 Sol’s long-task estimate ranged from 11.3 to more than 270 hours depending on scoring rules.
Anthropic's Claude Fable 5.1 reaches 90% ARC-AGI-2 coverage at 32% lower cost per task, with major gains on science and terminal benchmarks.