Cerebras adds Qwen 3.8 27B at ~1500 tokens/s

Original: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

Why This Matters

High-speed public inference for a 27B-parameter model lowers barriers for developers building latency-sensitive AI applications.

Cerebras Inference has added Qwen 3.8 27B (27 billion parameters) to its public model catalog, offering inference speeds of approximately 1,500 tokens per second with a 64k (free) / 128k (paid) context window.

Cerebras has listed Qwen 3.8 27B on its public inference endpoints, alongside the existing OpenAI GPT OSS 120B model. The Qwen 3.8 27B model runs at approximately 1,500 tokens per second, while the 120B model reaches around 3,000 tokens per second. Both models are available on the free trial and pay-as-you-go tiers, subject to rate limits. Cerebras states that all models on its public endpoints are original, unpruned versions. The company uses selective weight-only quantization for storage — storing weights in partial 16-bit, 8-bit, or 4-bit formats — while keeping activations, attention, and KV cache in full precision. Cerebras also notes that its REAP (Router-weighted Expert Activation Pruning) research models are available on Hugging Face for experimentation but are not served through the production API. Additional model families, dedicated capacity, and higher throughput are available through Dedicated Endpoints.

Source

inference-docs.cerebras.ai — Read original →