Jeff: 0.8B zero-shot decision models, ~30ms per call

Original: Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

Why This Matters

Tiny, fast, fine-tunable classifiers lower the barrier for on-device AI decision-making without API costs.

Developer firelex released Jeff, a set of fine-tuned 0.8B models (Qwen3.5 and Gemma 4) for zero-shot classification. Targeting fast local inference, Jeff delivers ~22ms per decision on an RTX PRO 6000 and ~28ms on Apple M4 Max, using the same API format as the Jev framework.

Jeff is an open-source project on GitHub offering small, fine-tuned classification models built on Qwen3.5 and Gemma 4 at roughly 0.8 billion parameters. The core promise: describe a situation and a list of options in plain text, and the model returns calibrated probabilities for each option in a single forward pass — no generated text, no output parsing.

Because it is zero-shot, the option labels never need to appear in training data. The repo lists support queues, user intent detection, content moderation, voice commands, and game moves as target use cases. Performance benchmarks show Jeff approaching or occasionally beating Jev, a larger reference model, though the project is upfront that raw reasoning depth cannot match a bigger model.

Where Jeff shines is fine-tuning speed. The README documents a voice-navigation task where held-out accuracy jumped from 31.7% to 95.8% after less than 30 minutes of fine-tuning on a single GPU. The project was built entirely on local hardware. It uses the same request format as Jev, making it a drop-in alternative for latency-sensitive pipelines.

Source

github.com — Read original →