Kog: Software-only approach to unlock 30x faster GPU inference
Original: Kog is going deeper to squeeze more inference out of GPUs
Why This Matters
Software-based GPU inference acceleration could reduce AI infrastructure costs without requiring new hardware investment.
French AI startup Kog claims its Kog Inference Engine (KIE) can deliver 30x faster LLM inference on standard datacenter GPUs like AMD MI300X and Nvidia H200. After a May tech preview generated 200 business leads via Hacker News, the company is now scaling its approach from small 2B-parameter models to large LLMs.
French startup Kog is pursuing faster AI inference through pure software optimization rather than purpose-built silicon, a strategy that contrasts with Cerebras's IPO debut in May 2026. CEO Gaël Delalleau told TechCrunch that a May tech preview on Hacker News generated 200 tangible business leads after demonstrating 3,000 tokens per second (TPS) per request on standard datacenter GPUs. That demo used Laneformer 2B, a purpose-built 2-billion-parameter model now open sourced, but Kog's stated goal is 30x faster inference on full-scale LLMs. Delalleau argues that newer GPUs offer growing memory bandwidth that software can unlock, calling the notion that GPUs are poorly suited for decoding 'a misconception.' Initial target customers include software engineering teams frustrated by multi-hour Claude Code waits, and app/game generation platforms where faster outputs translate to more revenue. Kog found that prospective customers were unwilling to fine-tune small models, prompting the team to shift focus toward accelerating larger models. The company's approach is compared to Stanford's Hazy Research lab for its low-level GPU optimization depth. Its seed round was co-led by Varsity VC, a firm run by Delalleau's former co-founder from his 2009 startup Stribe.