Redis Creator Launches ds4: Local LLM Engine in C
Original: From the creator of Redis; run LLM locally with ds4
Why This Matters
Brings 284B-parameter frontier MoE inference to consumer and pro hardware without cloud dependency.
Salvatore Sanfilippo (antirez), creator of Redis, released DwarfStar 4 (ds4), a narrow C inference engine for running DeepSeek V4/V4.1, GLM 5.x, and Qwen3.8 locally on high-memory Mac, CUDA, and ROCm machines. MIT licensed.
DwarfStar 4 (ds4) is a purpose-built, open-source inference engine written in C that runs frontier open-weight models locally. It supports DeepSeek V4 and V4.1 Flash (284-billion-parameter mixture-of-experts), GLM 5.x, and Qwen3.8 Flash — but only specific GGUF layouts, not generic files. This is a deliberate design choice: ds4 validates each supported model end-to-end rather than attempting broad compatibility.
The engine uses asymmetric 2-bit quantization to compress routed experts while preserving critical shared model paths — making these large MoE models practical on machines like Apple Silicon Macs with 64GB+ RAM, NVIDIA DGX Spark hardware, or AMD Strix Halo ROCm setups. One notable engineering detail: the KV cache is keyed by SHA1 of the prompt prefix and persisted to SSD, so server restarts don't require full re-prefill.
ds4 ships three interfaces from the same model state: a CLI (`./ds4`), an OpenAI- and Anthropic-compatible HTTP server (`./ds4-server`), and a native coding agent (`./ds4-agent`). Additional features include tensor parallelism, session batching, speculative decoding, and vision input support. Setup is three steps: clone the repo, download a model GGUF via script, and build for the target backend.