Bonsai 2 (Qwen 3.8 27B) as an OpenClaw Fallback: 8GB, 85 tok/s on a 5070
PrismML's Ternary-Bonsai-2-27B is a quantized Qwen3 3.8 27B build that a r/openclaw user is running locally as an OpenClaw fallback model when their ChatGPT and Grok quotas run dry. Their claim: ~8GB footprint, 98% of the base Qwen3 27B's quality retained, and text plus vision support.
Numbers from the setup
- Output speed: 85 tok/s
- Hardware: RTX 5070 with 16GB VRAM, 64GB system memory
- Disk footprint: ~8GB
- Quality retention: reported 98% of Qwen3 3.8 27B
The two GGUFs you need
Files come from the prism-ml/Ternary-Bonsai-2-27B-gguf repo on Hugging Face:
Ternary-Bonsai-2-27B-PQ2_0.gguf (text) Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf (vision)
Both are required if you want the multimodal path — the mmproj file is the vision projector, the PQ2_0 file is the ternarized text weights.
Setup notes
Bonsai 2 requires PrismML's own fork of llama.cpp — it will not load in stock llama.cpp. The author had GPT wire up the fork on a separate AI machine from the OpenClaw host, with the model configured to:
- Dynamically load when OpenClaw calls it
- Unload after 10 minutes of inactivity
- Act as the fallback when GPT 5.6 and Grok hit usage limits
What it actually did
The interesting part isn't the tok/s — it's the agentic behavior. Per the post, it handled web search via a browser, VM control, installing and using new software, and running existing workflows, to the point the author says they "can't tell it's not a frontier model."
That's a single user's report, not a benchmark, so treat the quality claims as anecdotal. The load/unload-on-idle pattern is the practical takeaway: it keeps a local 27B resident only when your paid provider is rate-limited, which is a sensible way to avoid permanently burning VRAM on a box that's also doing other work.
If you have a GPU with 16GB+ VRAM, the PQ2_0 + mmproj pair is small enough to be worth testing against your own agentic tasks before deciding whether it holds up for you.
📖 Read the full source: r/openclaw
👀 See Also

Claudius: Open-Source Embeddable AI Chat Widget for Claude
Claudius is an open-source, self-hosted chat widget powered by Claude that can be embedded on any website with one script tag. It runs on Cloudflare Workers with a React frontend and includes features like custom system prompts, rate limiting, and accessibility compliance.

WhatsApp AI Assistant Built with Claude Code as OpenClaw Alternative
A developer built a WhatsApp AI assistant using Claude Code as the agentic brain, with a local relay server for WhatsApp webhooks and MCP server bridging. The project includes Arcade for scoped auth to Google Calendar, Gmail, and Slack.

Claude Session Tracker: Auto-Save Claude Code Sessions to GitHub Issues
A new tool called claude-session-tracker automatically saves Claude Code sessions to GitHub Issues, logging every prompt and response as comments with timestamps. It creates one GitHub Issue per session linked to a Projects board and works through Claude Code's native hook system without consuming context tokens.

Ultimate Unreal Engine MCP: Claude Code Can Now Build and Verify Unreal Engine Levels with 132 Tools
Open-source MCP server exposes 132 tools across 26 domains, letting Claude spawn actors, set UPROPERTY values, take viewport screenshots, navigate cameras, and self-correct after mutations.