AI Guardrails Are Blocking Legitimate Cybersecurity Researchers
Original: How AI guardrails are impeding the work of offensive cybersecurity researchers
Why This Matters
AI security restrictions risk undermining the defensive cybersecurity ecosystem by limiting legitimate vulnerability research tools.
AI safety restrictions from Anthropic and OpenAI are hampering offensive cybersecurity researchers who use models to probe vulnerabilities. Experts warn that broad guardrails prevent legitimate security work, including exploit validation and vulnerability research, while Anthropic's Mythos models faced U.S. export controls in June 2026.
AI companies including Anthropic and OpenAI have implemented strict guardrails and vetted access programs to prevent misuse of their models by malicious actors. However, these restrictions are increasingly blocking legitimate offensive cybersecurity researchers whose job involves probing systems for weaknesses before criminals do.
In June 2026, the U.S. government imposed export control restrictions on Anthropic's Mythos and Fable models, partly prompted by reports of guardrail bypass vulnerabilities. The controls on Fable 5 were lifted on July 1, while Mythos 5 has been reintroduced only to vetted U.S. organizations. Both Anthropic and OpenAI offer formal vetting programs — the Cyber Verification Program and Trusted Access for Cyber program, respectively — granting approved researchers fewer restrictions.
Critics argue the guardrails are overly broad. Security researcher Mark Dowd, a longtime zero-day vulnerability researcher who sells exploits to Western governments, stated he is uncomfortable with 'large companies making arbitrary decisions about what is safe in security.' Chris Anley, Chief Scientist at NCC Group, noted that asking an AI to attempt to exploit a bug is essential to confirming whether a vulnerability is real and worth fixing — a step guardrails can block entirely, harming defenders in the process.