← Back

LLMs Turn Against LLMs: Automated Exploits Shift AI Security

Sep 19, 2026
LLMs Turn Against LLMs: Automated Exploits Shift AI Security

'''Independent security researchers from Hacktron AI have successfully used Anthropic's Claude to automate the discovery of vulnerabilities in OpenAI's ChatGPT, marking a pivotal escalation in AI security threats. This development moves beyond theoretical "jailbreaking" into automated, cross-platform exploits, demonstrating that foundation models themselves can be weaponized to attack rivals. It fundamentally shifts the AI security paradigm from defending against human red-teaming to defending against AI-driven attackers, a far more scalable threat. This mirrors the recent rise of AI-powered code generation tools, but applied to offensive security, creating an urgent need for a new class of automated, AI-native defense mechanisms. This "LLM-on-LLM" attack vector creates an asymmetric advantage for attackers, who can now run continuous, low-cost automated probes against high-value commercial models. The primary losers are proprietary model providers like OpenAI and Google, whose closed-source nature now appears more brittle, as they cannot leverage public bug bounty programs as effectively against these new threats. Winners include specialized AI security firms (such as Adversa AI or HiddenLayer) and, ironically, the providers of the attack models like Anthropic, whose tool