Stacked Defense Layers Drop Prompt Injection to 0 in Claude Code

Anthropic is claiming that stacked defense layers can drop prompt injection attacks to zero, even against unseen attacks. Boris Cherny, in a discussion about Claude Code's new Auto Mode, outlined the layered approach: model training, intent classifier, and input probes. He also confirmed the classifier is now free, responding to token cost concerns with: "We are making the classifier free. You should not need to pay for safety."
How the Layers Stack Up
- Model training: The base model is fine-tuned to recognize and reject injection patterns.
- Intent classifier: A separate classifier checks the user's intent behind each prompt, filtering out malicious requests before execution.
- Input probes: Active probes test inputs for known injection vectors, adding another barrier.
According to the source, this combination reduces prompt injection rates to 0% on unseen attacks. The key is that no single layer is perfect, but stacked they catch the vast majority of attempts.
Classifer Now Free
Boris Cherny's quote is direct: "We are making the classifier free. You should not need to pay for safety." This addresses a common pain point—token cost—for developers who worried that adding safety checks would drain their budgets. If you've been bypassing safety features to save tokens, that's no longer a concern.
This is particularly relevant for teams building autonomous agents with Claude Code, where prompt injection is a real risk—malicious prompts in web content or tool outputs could hijack the agent. With the classifier free, you can enable it without worrying about extra fees.
The full blog post is about Auto Mode in Claude Code, but this security detail stands out. For developers, it's a signal that Anthropic is treating safety as a baseline feature, not a premium add-on.
📖 Read the full source: r/ClaudeAI
👀 See Also
OpenClaw 2026.9.2 Prompt Injection Attempt: How It Happened and What to Learn
An attacker sent a prompt-injection payload to an OpenClaw WhatsApp channel, but the agent's own self-detection probe exposed it. No damage was done—minus a few read-only greps. What can we learn about structural trust boundaries?

Cisco source code stolen via Trivy supply chain attack
Cisco's internal development environment was breached using stolen credentials from the Trivy supply chain attack, resulting in the theft of source code from over 300 GitHub repositories including AI-powered products and customer code.

Threat data from 91K AI agent interactions: Tool abuse up 6.4%, new multimodal attacks
Analysis of 91,284 AI agent interactions from February 2026 shows tool/command abuse increased 6.4% to 14.5%, with tool chain escalation as the dominant pattern. RAG poisoning shifted to metadata attacks (12.0%), and multimodal injection via images/PDFs emerged at 2.3%.

Anthropic reveals industrial-scale Claude AI data extraction by Chinese labs
Anthropic confirmed Chinese AI labs used over 24,000 fraudulent accounts to scrape 16 million exchanges from Claude, extracting safety guardrails and logic structures for military and surveillance systems.