Running Claude Code Offline on an M3 Pro with Qwen3.6: 4 Fixes That Made It Work

Claude Code connects to a local model on an Apple M3 Pro (18 GPU cores, 36GiB unified memory, ~150 GB/s bandwidth) running qwen3.6:35b-a3b-coding-nvfp4 — a 35.1B-parameter MoE model with ~3B active per token, NVFP4 quantized, ~21GB on disk and ~20GiB resident. The setup took a Kubernetes incident from investigation to PR: found root cause, wrote patch, pushed branch, filed PR via gh — all air-gapped. Four fixes turned a model that timed out in 10 minutes into one that closes the loop. Speed is hardware-bound; capability is not.
Stack and Environment
- Hardware: Apple M3 Pro, 18 GPU cores, 36 GiB unified memory, ~150 GB/s memory bandwidth
- Model:
qwen3.6:35b-a3b-coding-nvfp4 - Runtime: Ollama 0.24.0, MLX runner (Apple Silicon-native)
- Client: Claude Code v2.1.84 pointed at local Ollama endpoint
Key environment variables (set in a launchd plist for persistence):
OLLAMA_MLX=1
OLLAMA_CONTEXT_LENGTH=32768
OLLAMA_FLASH_ATTENTION=1
OLLAMA_MULTIUSER_CACHE=1
OLLAMA_KEEP_ALIVE=24h
OLLAMA_NO_CLOUD=1
Setup Steps
- Install Ollama 0.24.0+
ollama pull qwen3.6:35b-a3b-coding-nvfp4(~21GB one-time)- Start server with the env vars above
- Launch Claude Code:
ANTHROPIC_BASE_URL=http://localhost:11434 MAX_THINKING_TOKENS=0 claude --model qwen3.6:35b-a3b-coding-nvfp4 - Smoke test:
Run kubectl get pods -A and tell me if anything appears unhealthy
Performance Notes
First tool call: seconds (thinking disabled). Prefill (loading ~25K tokens) takes ~60s. Subsequent turns are faster due to prefix caching (OLLAMA_MULTIUSER_CACHE). The model stays loaded via OLLAMA_KEEP_ALIVE=24h. Burst of 404s in Ollama log during prefill is normal (fix #4).
The MoE architecture is key: only ~3B active per token, so runtime cost resembles a 14B dense model while answers approach 35B. A dense 35B doesn't fit 36GiB.
📖 Read the full source: HN LLM Tools
👀 See Also

Claude Workflow Library: 10 Complete AI Workflows for Non-Technical Users
A free GitHub repository provides 10 complete AI workflows for Claude users without technical backgrounds, including study, research, writing, business, content creation, decision making, learning, job search, productivity, and life planning systems.

Pepper MCP Server for iOS Simulator Interaction and Debugging
Pepper is an MCP server that injects a dylib into iOS simulator apps via DYLD_INSERT_LIBRARIES, enabling real-time interaction, screen reading, button tapping, variable inspection, and network traffic monitoring through a WebSocket bridge.

Soul MCP Server Adds Persistent Memory and Safety for Local LLMs
Soul is an open-source MCP server that provides persistent memory across sessions for local LLMs with two commands: n2_boot at start and n2_work_end at end. It includes Ark safety features that block dangerous commands like rm -rf and DROP DATABASE at zero token cost, plus cloud storage configuration.

Agent Wake Skill for OpenClaw: Notify Discord When Tasks Complete
A developer created agent-wake.py, a Python script that Claude Code calls after tasks finish. It sends Discord pings and fires wake events via the gateway HTTP API, prompting the agent to post summaries automatically.