GPT-6 Astra: Looped Transformers and Hidden Reasoning Explained
Original: GPT-6 Astra, looped transformers, and hidden reasoning
Why This Matters
GPT-6 Astra's architecture and benchmark leadership signal a significant shift in frontier model design and capability.
OpenAI released GPT-6 Astra, its latest flagship model, achieving 99.9% on ARC-AGI-3 (vs. 7.8% for GPT-5.6) and topping coding/math benchmarks. The model features looped transformer architecture and reportedly conceals its chain-of-thought reasoning trace.
OpenAI launched GPT-6 Astra to widespread attention, with AI researcher Sebastian Raschka offering an in-depth analysis of its performance, architecture, and implications. According to Raschka, Astra is 'the best model I've used as of this writing,' excelling across writing, math, coding, and especially 3D rendering and animation tasks. On the ARC-AGI-3 benchmark, Astra scored 99.9%, compared to just 7.8% for its predecessor GPT-5.6. On the independent Artificial Analysis Intelligence Index v4.2 and Coding Agent Index v1.4, Astra leads the field, though its margin over competitors is not as dramatic as in some self-reported benchmarks. Raschka notes that independent benchmarks like those from Artificial Analysis use shared harnesses (e.g., Stirrup, Terminus 2) for more apples-to-apples comparisons, but that model training pipelines are often optimized for specific harnesses, which may understate Astra's real-world advantage. Architecturally, Astra is rumored to use 'looped transformers' (also called recurrent depth), where transformer blocks are reused across multiple passes rather than stacked linearly. Raschka explores how this design relates to rumors that Astra hides its internal chain-of-thought reasoning, a topic he ties to recent research on looped transformer behavior.