Frontier AI Models Alarmingly Vulnerable to Jailbreaks, FAR.AI Reports

Original: It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

Why This Matters

Demonstrates that frontier AI safety guardrails remain inconsistent and cheaply bypassable, highlighting the urgency of external AI regulation.

AI safety nonprofit FAR.AI tested safety guardrails of models from Anthropic, OpenAI, Google, and SpaceXAI using an automated jailbreak tool. Grok showed the most vulnerabilities (448 jailbreaks found), followed by Gemini (249), while Claude, Fable, and GPT resisted all automated attacks.

California-based AI safety nonprofit FAR.AI released a report testing the jailbreak resistance of frontier AI models from four major US companies: Anthropic's Claude Opus 4.8 and Fable 5, OpenAI's GPT 5.5 and 5.6, Google's Gemini 3.1 Pro, and Elon Musk's SpaceXAI's Grok 4.3 and 4.5. The organization's automated tool generates over 1,000 prompt variations designed to bypass safety guardrails, targeting harmful outputs such as cyberattack planning, software exploits, and chemical or biological weapon development details.

Grok proved most susceptible with 448 successful jailbreaks at a cost of just $58, while Gemini yielded 249 jailbreaks for $278. Claude, Fable, and GPT resisted all automated attacks, though FAR.AI noted these models may still be vulnerable to more sophisticated, complex jailbreak methods.

FAR.AI CEO Adam Gleave stated: 'AI models right now are less regulated than restaurants,' calling voluntary self-regulation 'nonsense' and advocating for externally imposed standards. Google DeepMind's AGI Safety Director Rohin Shah cautioned against treating the report as a comprehensive safety assessment, noting ongoing red-teaming and multi-layered protections. Anthropic and OpenAI also issued statements emphasizing their continued investment in evolving safety systems.

Source

wired.com — Read original →