mistral.rs Adds Support for Gemma 4 12B: Multimodal, Agentic, and MTP

mistral.rs now supports Gemma 4 12B with multimodal, agentic, and Multi-Turn Prediction (MTP) features. This release includes web search and sandboxed code execution for building agentic apps, plus audio, image, and video input.
Installation
Single-line install for Linux/macOS and Windows:
# Linux/macOS
curl --proto '=https' --tlsv1.2 -sSf https://raw.githubusercontent.com/EricLBuehler/mistral.rs/master/install.sh | sh
Windows
irm https://raw.githubusercontent.com/EricLBuehler/mistral.rs/master/install.ps1 | iex
Running with Agent & Quantization
Launch an OpenAI- and Anthropic-compatible HTTP server with a built-in web UI at localhost:1234/ui:
mistralrs run --agent -m google/gemma-4-12B-it --quant 4Enabling MTP (Multi-Turn Prediction)
To use MTP, add the --mtp-model flag with the assistant model:
mistralrs run --agent -m google/gemma-4-12B-it --quant 4 --mtp-model google/gemma-4-12B-it-assistantKey Features
- Full multimodal support: audio, image, and video
- Web search and sandboxed code execution for agentic workflows
- OpenAI and Anthropic-compatible HTTP server
- Built-in web chat UI at
localhost:1234/ui
For more details: GitHub | Documentation
📖 Read the full source: r/LocalLLaMA
👀 See Also

Prism MCP v5.1 adds 10x memory compression and agent learning from corrections
Prism MCP v5.1 introduces 10x memory compression via TurboQuant ported to TypeScript, enabling millions of memories on a laptop without vector databases. The update adds agent learning from user corrections and a visual knowledge graph interface.

Reddit User Tests Hermes AI Agent's Self-Learning Feature, Finds Critical Flaws
A Reddit user tested Hermes AI agent's self-learning feature, which automatically creates skills from markdown files. The user found it always evaluates its own results as successful, even when output is incorrect, and overwrites manual edits.

Using /probe to catch AI hallucinations before writing code
A developer shares a technique called /probe that forces AI-generated plans to make numbered claims with expected values, then probes the real system to catch discrepancies. The method caught four factual errors in Claude's description of its own JSONL format that would have caused code bugs.

ModelFitAI: Deploy AI Agents Without VPS Setup, Built with Claude Code
ModelFitAI is a platform that lets developers deploy AI agents directly on its infrastructure, eliminating VPS setup, Docker configuration, and SSH sessions. The entire platform was built using Claude Code by a solo founder.