Run 125B AI Model on RTX 4090 at 100 tok/s
Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Why This Matters
Brings frontier-scale 125B model inference to $1,500 consumer GPUs, no cloud required.
Open-source project Strata lets consumer GPUs (12GB+ VRAM) run Qwen3.8-Flash-Next, a 125B-parameter model, at ~100 tokens/sec on a single RTX 4090, with one-click Windows/Linux install and local OpenAI/Anthropic-compatible API.
Strata is an open-source inference engine that brings Alibaba's 125-billion-parameter Qwen3.8-Flash-Next model to consumer hardware — specifically NVIDIA or AMD GPUs with at least 12 GB of VRAM. The project, posted on GitHub by developer Niko1221, has already accumulated 10.2k stars and 906 forks, signaling strong community interest.
The engine reportedly hits around 100 tokens per second on an RTX 4090, a speed that makes real-time chat and code generation practical on desktop hardware. Setup is a single batch file on Windows or a shell script on Linux. Once running, Strata exposes a local server with OpenAI- and Anthropic-compatible API endpoints, meaning existing tools and applications can connect without modification.
Features include image input support, 128K context windows, and quantization options (e.g., IQ3_S) that allow the 125B model to fit within consumer VRAM budgets. The README shows a voxel art image generated from a single prompt on an RTX 5070 as a demonstration of multimodal capability. The project uses ggml as its tensor backend and ships with Docker support for more controlled deployments.