Cerebras & OpenAI Launch GPT-5.6 Sol Ultrafast at 750 Tokens/sec
Original: Accelerating GPT-5.6 Sol Ultrafast
Why This Matters
Ultra-high-speed frontier inference removes the longstanding speed-vs-intelligence tradeoff for enterprise AI deployments.
Cerebras and OpenAI announced Ultrafast Mode on August 13, 2026 — a new API service tier powered by Cerebras hardware that delivers GPT-5.6 Sol at up to 750 output tokens per second, with no quality degradation, initially available to select customers.
Cerebras and OpenAI have revealed Ultrafast Mode, a new service tier for the OpenAI API powered by Cerebras infrastructure. The offering runs GPT-5.6 Sol — OpenAI's flagship model for legal, financial, and engineering workloads — at up to 750 output tokens per second. According to Cerebras, this makes it 11x faster than Claude Fable 5 and 5x faster than Opus 4.8 on Fast mode, based on speeds reported by Artificial Analysis. In a head-to-head benchmark using Humanity's Last Exam (HLE), a 2,500-question test designed for PhD-level knowledge across chemistry, economics, and literature, GPT-5.6 Sol Ultrafast completed all questions in 11 hours and 11 minutes, compared to 78 hours and 27 minutes for Claude Fable 5 — roughly 7x faster at comparable accuracy. On GDP-Val, a benchmark for economically valuable knowledge work, Ultrafast delivered a 5.6x end-to-end speedup with no quality loss. OpenAI's Rohan Varma stated: 'With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate.' Targeted use cases include production incident response, cybersecurity threat detection, and real-time agentic workflows. Access is initially limited to select customers, with broader availability planned over time.