Merlin Research releases Qwen3.5-4B-Safety-Thinking model for structured reasoning

Merlin Research has released Qwen3.5-4B-Safety-Thinking, a 4 billion parameter safety-aligned reasoning model built on Qwen3.5. This model is specifically designed for structured 'thinking' and safety applications in real-world scenarios, with particular focus on agent systems.
Key improvements and features
- Improved ability to accurately follow strict instructions in prompts
- Based on the use of Bloom and Petri methods from Anthropic
- Resistant to hacking attempts
- Increased resistance to 'abnormal' and adversarial prompts
- Up to 1 million token context window
- Uses frameworks from Anthropic - Bloom and Petri
The model is available on Hugging Face at MerlinSafety/Qwen3.5-4B-Safety-Thinking.
For developers working with AI agents, this model represents a specialized tool for safety-critical applications where structured reasoning and resistance to prompt manipulation are priorities. The integration of Anthropic's Bloom and Petri methods suggests a focus on constitutional AI approaches to alignment.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Andon Labs' AI Agent Mona Runs a Real Cafe in Stockholm — Full Breakdown
Andon Labs gave an AI agent named Mona a lease and real money to open a cafe in Stockholm. She handled bureaucracy, suppliers, and hiring, but hit walls like BankID and had to make suboptimal choices.

Micron's $200B Investment Aimed at AI Memory Constraints
Micron commits $200 billion towards addressing AI memory bottlenecks, aiming to enhance AI processing capabilities.

Pope Leo XIV's 'Magnifica Humanitas': A 40,000-Word Encyclical on AI Disarmament
Pope Leo XIV releases Magnifica Humanitas, a 40,000-word encyclical calling for AI disarmament, critiquing autonomous weapons, data colonialism, and tech monopolies. Co-founder of Anthropic present at release.

AI's Brokenomics: Anthropic's Mythos/Fable Export Ban Chaos
Anthropic's 'too dangerous to release' Mythos model was jailbroken within days, leading to US export controls banning non-US citizen access. Fable's guardrails failed when Amazon researchers broke them, triggering a national security rollback.