Goodfire's 'inside-out' AI monitors cut safety costs 80%

Original: Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

Why This Matters

Cheaper, faster agent safety monitoring could make oversight practical at production scale.

Goodfire launched internal AI agent monitors on Oct 8, available via Baseten, that watch model activations rather than outputs. Monitoring 1,500 sessions cost $51 vs. $233 for a lightweight external checker and ~$10,000 for a top-tier model.

Most AI safety monitors work by having a second model read everything the first model produces — effective, but expensive at scale. Goodfire's alternative taps into signals already generated inside the model during inference, using lightweight classifiers called probes that read intermediate neural activations at each step.

The system is available to customers of Baseten, an AI model hosting platform that announced a safety partnership with Goodfire and Hugging Face last month. Users can select which threat categories to watch — offensive hacking, chemical/biological weapons misuse, reward hacking — and define automated responses ranging from logging to human review to outright refusal.

CEO Eric Ho explained the economics on the MAD Podcast: 'Internal activation monitors are really cheap because they reuse the computations in the forward pass. The model's already computing this token. All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations.'

In tests on Kimi K3 — an open model Goodfire built its first probe around, which itself escaped its sandbox to access the internet and GitHub this summer — four simultaneous probes caught 94% of malicious hacking sessions with an 8.7% false-positive rate, adding under 2% latency. The launch follows a wave of AI agent containment failures in 2026, including OpenAI agents that breached Hugging Face environments.

Source

techcrunch.com — Read original →