OpenAI launches 'Ultrafast' mode for GPT-5.6 Sol at 14x speed
Original: OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed
Why This Matters
Ultra-high inference speed signals a new competitive front in enterprise AI deployment.
OpenAI has released a new mode called Ultrafast for its GPT-5.6 Sol model, delivering up to 750 output tokens per second — 14x faster than standard processing. Powered by Cerebras chips, the preview is currently limited to select customers.
OpenAI has introduced 'Ultrafast,' a new inference mode for its latest model, GPT-5.6 Sol, capable of generating up to 750 output tokens per second — 14 times the speed of standard processing. The feature is powered by OpenAI's existing partnership with chip maker Cerebras. In a blog post published Thursday, OpenAI stated: 'Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.' The company positions Ultrafast as suitable for enterprise use cases including incident response, customer service and support, financial market analysis, and e-commerce. The mode is currently available in preview to a limited set of customers, with OpenAI saying it will expand access as 'capacity grows.' Competitor Anthropic also offers a fast mode for its Claude models, though OpenAI claims its speed figures surpass what Claude's fast mode currently delivers.