Researchers Extract Hidden 'Reasoning Traces' from Claude, GPT, Gemini

Original: A New Trick Reveals AI Models’ Inner Thoughts

Why This Matters

Exposes a systemic vulnerability in frontier AI model APIs with implications for IP protection and AI security.

Computer scientists at University of Tübingen developed a method to extract hidden reasoning traces from frontier AI models via API. Findings suggest Chinese model Kimi K3 closely mirrors Claude Opus 4.8 and GPT 5.6 Sol reasoning patterns, raising distillation concerns, though causal proof remains absent.

Researchers from the University of Tübingen, Max Planck Institute, MATS Research, and Snyk identified a shared vulnerability across frontier AI models from OpenAI, Anthropic, and Google that allows extraction of their hidden 'reasoning traces'—the internal chain-of-thought steps models use to solve complex problems. Lead researcher Alexander Panfilov stated: 'All major frontier model providers we tested share this vulnerability. It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks.' The team demonstrated the method could also recover personal data such as passwords and API keys from a model's inner reasoning, though that specific vulnerability has since been patched. Most notably, the researchers found that Moonshot AI's open-weight model Kimi K3 produces reasoning outputs strikingly similar to those of Claude Opus 4.8 and GPT 5.6 Sol for certain prompts. The paper stops short of proving distillation, stating it 'cannot causally establish distillation.' By contrast, DeepSeek and Thinking Machines' Inkling did not show similar patterns to Claude Opus. Distillation—copying capabilities from existing models into new ones—has become increasingly controversial. In February, OpenAI alleged DeepSeek copied its model to build R1; in June, Anthropic told US lawmakers that Alibaba systematically distilled its models to produce Qwen. Moonshot AI and Z.ai did not respond to requests for comment.

Source

wired.com — Read original →