Build Your Own Single-Pass Decision Model

Original: Build your own decision model

Why This Matters

Single-pass constrained decoding slashes inference cost for classification-style tasks.

Developer Nish Tahir explains how to build a decision model using constrained token decoding with Qwen3-1.7B, replacing multi-pass generation with a single forward pass over fixed answer options.

Standard LLMs generate structured output token-by-token—producing a JSON response might take 11 separate passes. Tahir's post contrasts this with decision models like Jev, which assume a closed set of possible outputs and answer in a single forward pass by masking all vocabulary tokens except the allowed options (e.g., A through E), then picking the highest-probability token.

The tutorial walks through implementing this with Qwen/Qwen3-1.7B via Hugging Face Transformers. The key step: extract logits only for the target option token IDs, apply softmax over that constrained subset, and return the argmax. No beam search, no autoregressive loop.

Tested on a simple sky-color question, the model assigns 99.88% probability to 'Blue.' Evaluated on a random holdout from CommonsenseQA, F1 scores range from ~0.59 to ~0.64 across answer classes—reasonable for a 1.7B model with no task-specific fine-tuning. Tahir notes an important caveat: the output probabilities reflect next-token confidence, not calibrated correctness probability, so treating them as accuracy scores without additional training is misleading.

Source

nishtahir.com — Read original →