Inception Labs Releases Mercury 2.5 Diffusion LLM
Original: Mercury 2.5
Why This Matters
Mercury 2.5 demonstrates diffusion-based LLMs maturing into production-grade, latency-sensitive enterprise deployments.
Inception Labs launched Mercury 2.5, its most capable diffusion language model to date, claiming 40% intelligence gain over Mercury 2, 1,107 tokens/sec throughput, 260K token context, and pricing at $0.20/$0.75 per million tokens (80% off at launch).
Inception Labs released Mercury 2.5, described as the largest diffusion language model (dLLM) ever trained. The model delivers a 40% intelligence improvement over Mercury 2 while maintaining the same low-latency, low-cost profile. Performance is benchmarked as comparable to cost-optimized frontier models including GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Technical specs include 1,107 tokens per second on standard NVIDIA GPUs, a 260K token context window, and standard pricing of $0.20 per million input tokens and $0.75 per million output tokens — reduced 80% to $0.04/$0.15 at launch. Mercury 2.5 supports tunable reasoning, parallel tool calls, and schema-aligned JSON. Production use cases highlighted include search/RAG pipelines, voice agents, and coding assistants. OpenCall reported P99 response latency dropping from several minutes to one second after switching to Mercury. Augment Code reduced context compaction latency by 82% (150s to 27s) and cut costs by 90%. Inception Labs also announced previews of Mercury Voice, targeting under-170ms time-to-first-token for voice agents, and Mercury Router.