PinchBench Ranks Qwen and Nemotron on Mac Studio M3 Ultra: Nemotron Super Hits 99.2% Coding

✍️ OpenClawRadar📅 Published: September 24, 2026🔗 Source
Ad

PinchBench is an OpenClaw local model benchmarking tool. A developer posted results from running six models on a Mac Studio M3 Ultra (256GB RAM, 60-core GPU), with temperature at 0.8 and the context window set to 200,000 tokens where supported. The ranking below is lifted directly from that post — including the gaps you'd care about before picking a daily driver.

Full results

#ModelDisk sizeMax token windowTotal / Coding
1Qwen3.6 35B A3B20.4 GB262k87.9% / 92.6%
2QWEN3-Coder-next-Q884.8 GB262k85.5% / 97.9%
3Qwen3.5-122B-a10b-uncensored-hauhaucs-aggressive78.7 GB262k83.5% / 91.7%
4Nemotron-3-super86.1 GB1.05M78.2% / 99.2%
5QWEN3.8-flash-next95.43 GB262k71.2% / 79.2%
6Nemotron-3-nano-30b-a3b-mlx33.6 GB262k65.8% / 65.1%

The post also lists a Dolphin-Mistral-24b-venice-edition-mlx-8b at 25.1 GB with a 131k token window, but the benchmark total is truncated in the source text. Results for it aren't usable as written.

Ad

What stands out

  • Qwen3.6 35B A3B wins on balance. Highest total score (87.9%) and a strong 92.6% coding score, at only 20.4 GB on disk. It's also the author's daily driver.
  • QWEN3-Coder-next-Q8 is the coding specialist. 97.9% coding but a lower 85.5% overall — and 84.8 GB, over four times the disk footprint of the 35B A3B. The author had not used it before this test and says they'll try it.
  • Nemotron-3-super has the highest coding score at 99.2%, but its 78.2% total is the lowest of the top four. It also has the largest context window in the field at 1.05M tokens — 4x the Qwen models.
  • The Nemotron Super vs Nano gap is large. 99.2% vs 65.1% coding, 78.2% vs 65.8% total. Nano is less than half the size (33.6 GB vs 86.1 GB), which explains part of it.
  • QWEN3.8-flash-next underperformed for its size. At 95.43 GB — the largest model tested — it scored 71.2% / 79.2%, below Qwen3.5-122B (78.7 GB).

Two caveats from the author

First, they say they haven't found effective settings for QWEN3.8 on Mac and are asking for settings that work. The low score for QWEN3.8-flash-next shouldn't be read as a model limit until that's resolved.

Second, all numbers are single-machine, single-run at temperature 0.8 on M3 Ultra hardware. Treat the ranking as a starting point for your own PinchBench run, not a verdict — especially before spending 85-95 GB of disk on a download.

Who this is for

Anyone running local models in OpenClaw or similar agents on Apple Silicon with enough unified memory (64GB+) to load 30-120B models. If you're on 32GB or less, the 20.4 GB Qwen3.6 35B A3B is the only realistic pick from this list — which happens to be the top scorer anyway.

📖 Read the full source: r/openclaw

Ad

👀 See Also