PinchBench Ranks Qwen and Nemotron on Mac Studio M3 Ultra: Nemotron Super Hits 99.2% Coding
PinchBench is an OpenClaw local model benchmarking tool. A developer posted results from running six models on a Mac Studio M3 Ultra (256GB RAM, 60-core GPU), with temperature at 0.8 and the context window set to 200,000 tokens where supported. The ranking below is lifted directly from that post — including the gaps you'd care about before picking a daily driver.
Full results
| # | Model | Disk size | Max token window | Total / Coding |
|---|---|---|---|---|
| 1 | Qwen3.6 35B A3B | 20.4 GB | 262k | 87.9% / 92.6% |
| 2 | QWEN3-Coder-next-Q8 | 84.8 GB | 262k | 85.5% / 97.9% |
| 3 | Qwen3.5-122B-a10b-uncensored-hauhaucs-aggressive | 78.7 GB | 262k | 83.5% / 91.7% |
| 4 | Nemotron-3-super | 86.1 GB | 1.05M | 78.2% / 99.2% |
| 5 | QWEN3.8-flash-next | 95.43 GB | 262k | 71.2% / 79.2% |
| 6 | Nemotron-3-nano-30b-a3b-mlx | 33.6 GB | 262k | 65.8% / 65.1% |
The post also lists a Dolphin-Mistral-24b-venice-edition-mlx-8b at 25.1 GB with a 131k token window, but the benchmark total is truncated in the source text. Results for it aren't usable as written.
What stands out
- Qwen3.6 35B A3B wins on balance. Highest total score (87.9%) and a strong 92.6% coding score, at only 20.4 GB on disk. It's also the author's daily driver.
- QWEN3-Coder-next-Q8 is the coding specialist. 97.9% coding but a lower 85.5% overall — and 84.8 GB, over four times the disk footprint of the 35B A3B. The author had not used it before this test and says they'll try it.
- Nemotron-3-super has the highest coding score at 99.2%, but its 78.2% total is the lowest of the top four. It also has the largest context window in the field at 1.05M tokens — 4x the Qwen models.
- The Nemotron Super vs Nano gap is large. 99.2% vs 65.1% coding, 78.2% vs 65.8% total. Nano is less than half the size (33.6 GB vs 86.1 GB), which explains part of it.
- QWEN3.8-flash-next underperformed for its size. At 95.43 GB — the largest model tested — it scored 71.2% / 79.2%, below Qwen3.5-122B (78.7 GB).
Two caveats from the author
First, they say they haven't found effective settings for QWEN3.8 on Mac and are asking for settings that work. The low score for QWEN3.8-flash-next shouldn't be read as a model limit until that's resolved.
Second, all numbers are single-machine, single-run at temperature 0.8 on M3 Ultra hardware. Treat the ranking as a starting point for your own PinchBench run, not a verdict — especially before spending 85-95 GB of disk on a download.
Who this is for
Anyone running local models in OpenClaw or similar agents on Apple Silicon with enough unified memory (64GB+) to load 30-120B models. If you're on 32GB or less, the 20.4 GB Qwen3.6 35B A3B is the only realistic pick from this list — which happens to be the top scorer anyway.
📖 Read the full source: r/openclaw
👀 See Also

Open-Source Framework Uses Claude Code CLI for Automated GitHub Repo Monitoring
A developer has open-sourced a framework that runs Claude Code CLI on a cron schedule to triage GitHub activity across multiple repositories. The tool includes state tracking, deduplication, Discord notifications, and a pre-check system that avoids API costs when nothing has changed.

VibeAround: Local Daemon Connects Coding Agents to Telegram and Discord
VibeAround is a local daemon that connects coding agents like Claude Code, Gemini CLI, and Codex to IM platforms including Telegram and Discord. The tool features session handover with pickup codes to continue conversations across devices.

Time Complexity MCP: Static Analysis Tool Feeds Big-O Complexity to AI Coding Agents
Time Complexity MCP is an open-source MCP server that performs static code analysis to detect Big-O complexity, feeding the results directly to AI coding agents like Claude Code or Copilot without token consumption. It supports JavaScript, TypeScript, Python, Java, Kotlin, and Dart.

ClawNet: Peer-to-Peer AI Agent Network Without API Keys
ClawNet is a peer-to-peer network that allows AI agents to collaborate directly without API keys or platform fees. Installation is via a curl script, and features include a task bazaar, shell economy, and knowledge network.