Magnitude: Self-optimizing local inference engine for AI agents
Original: Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Why This Matters
Local inference speed is a hard limit on agentic AI usability; a 2x gain without new hardware matters.
YC S25-backed Magnitude has open-sourced a local inference engine that compiles and tunes its kernels per device, running open models up to 2x faster than llama.cpp. It supports Apple Silicon, NVIDIA, AMD, and CPU-only setups, with one-click integration for agents like Pi, OpenCode, and Codex.
Magnitude, a Y Combinator S25 company, has released an open-source inference engine on GitHub (5.7k stars, 397 forks at time of writing) designed specifically for agentic AI workloads. Unlike static runtimes, Magnitude compiles and tunes its compute kernels on the user's own hardware at setup time—yielding up to 2x faster throughput compared to llama.cpp on the same device. The engine works across Apple Silicon, NVIDIA GPUs, AMD GPUs, and CPU-only machines, aiming to cover the full range of developer hardware without separate builds. Integration is described as a single click for popular agents including Pi, OpenCode, Hermes, Codex, and others. Desktop builds are available for macOS, Windows, and Linux. The repo already shows multiple generations of the inference stack (inference-v2 through inference-v4), suggesting active iteration. No pricing has been announced; the project is Apache-2.0 licensed.