Opus 5.5 reasoning_extraction refusals: the trigger is often a word in your own prompt or tool schema
If your agents on Opus 5.5 (also Opus 5 and Fable 5.1) occasionally stop with stop_reason: "refusal" and category reasoning_extraction, the trigger is frequently in your own setup — not the user message. Anthropic's prompting guide and refusals page describe the category as covering requests that "push the model to reproduce its internal reasoning in the response text." Their documented fix is to remove "write out your reasoning" style instructions and read summarized thinking via display: "summarized" instead.
What the docs and billing tell you
- Server-side fallback does not retry this category — the refusal comes back to you.
- A refusal that arrives before any output is billed for this category (same for
bioandfrontier_llm).
The most useful evidence: FaultMaven #1751
They bisected a tool schema. A single string property called internal_reasoning was enough to trigger a refusal on opus-5-5, fable-5-1 and opus-5 — while opus-4-8 and sonnet-5 accepted the identical request. Renaming the property to evidence_trail and removing the word "reasoning" from the descriptions fixed it. The field name and the wording of what you ask for matter, not what the model would actually write.
Other reported triggers
- An LLM-judge prompt in sublang's playbook refused on every adjudication call.
- A plain
git commitin Claude Code (#97001). - A CHANGELOG update, reported in r/ClaudeCode.
- The poster's own OpenClaw setup: two refusals on Opus 5 (Sept 19 and 20), both in conversations about the agent's own pre-answer text showing up in chat — no explicit "reasoning" instruction, just discussing it.
Why a fallback chain matters
In the author's case, the first refusal was caught by OpenClaw's model fallback and answered by Sonnet 5. That is the practical takeaway: the gateway's fallback retries where Anthropic's server-side fallback will not. An agent whose fallback chain is only Opus 5.x / Fable simply fails.
Where to grep in a workspace
AGENTS.mdandSOUL.md- Skills
- Cron and heartbeat prompts
- Compaction or summary prompts
- Any JSON output format with a field like
"reasoning"or"thought_process"
Obvious candidates in text: "explain your reasoning", "show your thinking step by step", "include your chain of thought".
The open question from the thread: has anyone seen this fire on a prompt with none of that wording? That would say more about how wide the classifier actually is.
📖 Read the full source: r/openclaw
👀 See Also

Research shows personality affects Claude's self-correction, not Llama or Qwen
A researcher ran 23 experiments testing self-correction without guardrails across Claude, Llama, and Qwen. The main finding: personality profiles affect Claude's self-correction ability, with high directness catching all errors and low directness catching none. Llama and Qwen didn't self-correct even with identical prompts.

Liquid AI releases LFM2.5-350M model for agentic loops
Liquid AI released LFM2.5-350M, a 350M parameter model trained for reliable data extraction and tool use. It's under 500MB when quantized and outperforms larger models like Qwen3.5-0.8B in most benchmarks while being faster and more memory efficient.

Opus 4.6 Medium vs Low: Performance Differences and Pricing
Opus 4.6 medium costs approximately 50% more than the low version but addresses significant laziness issues found in the low-powered model. The medium version sits between low and high in performance benchmarks.
Microsoft Exec Called AI Scraping 'the Largest Theft of Labor in Human History,' Unredacted NYT v. OpenAI Filings Show
Newly unredacted filings in The New York Times' three-year copyright suit against OpenAI and Microsoft quote a Microsoft director calling scraping 'the largest theft of labor in human history.' Internal data also shows Copilot cut NYT click-throughs by up to 93%.