Local Qwen 3.6 vs Frontier Models on a Coding Primitive: Single-File HTML Canvas Driving Animation

A Reddit user ran a head-to-head comparison of local quantized models versus frontier web-based models on a specific coding primitive: generating a single HTML file with a full-page canvas animation of a side-view car driving with parallax scrolling, spinning wheels, and cinematic lighting.
The Prompt
The exact prompt asked for a single HTML file with no libraries, a full-page canvas, realistic side-view car animation, layered parallax scenery, spinning wheels, subtle body motion, smooth looping, and cohesive sky/lighting.
Models Tested
Frontier (web-based via Perplexity, tok/s not measured):
- Claude Sonnet 4.6 Thinking (used internet for reasoning)
- Gemini 3.1 Pro Thinking
- GPT 5.4 Thinking
- Kimi k2.6 Thinking
Local (Ryzen 5 5600, 24 GB DDR4-3200, RX 5700 XT 8GB):
- Qwen3.5 9B Q4_K_M — ~50 tok/s
- Qwen3.6-27B (Claude-opus-reasoning-distilled) Q4_K_M — 2.65 tok/s
- Qwen3.6-27B Q4_K_M — 2.70 tok/s
- Qwen3.6-31B A3B Q4_K_M — 12.13 tok/s
- Gemma-4-31b-it — 1.91 tok/s
- Qwen3.5 4B Q8 — 60 tok/s (used internet for reasoning)
- Qwen3.5 4B Q4_K_M — 80 tok/s (used internet for reasoning)
Results & Subjective Ranking
The ranking for this specific task:
- Kimi k2.6 Thinking — cleanest overall visual result
- Qwen3.6-27B Q4_K_M (local) — stronger than expected; good parallax and road feel
- Qwen3.6-27B Claude-opus-reasoning-distilled — close third
The local 27B quant delivered more natural motion and layering than some frontier outputs for this specific visual primitive. The poster noted they expected frontier models to outperform local quants more clearly.
The user only changed HTML <title> tags to track which model generated which file. Outputs are shared in the thread along with screenshots/GIFs of the running animations.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Claude Code v2.1.147: Pinned Sessions, /code-review, and Dozens of Fixes
Claude Code v2.1.147 introduces pinned background sessions, renames /simplify to /code-review with effort levels and --comment, plus fixes for PowerShell, MCP, Windows, and more.

Current LLM Cost Comparison: Deepseek, Qwen, MiniMax vs OpenAI
A Reddit analysis shows Deepseek-V3.2 at $0.26/$0.38 per million tokens is approximately 10x cheaper than GPT-4 while delivering GPT-5 class benchmark performance, with Qwen3.5 and MiniMax-M2.5 offering competitive alternatives to Claude and OpenAI.

Greg Kroah-Hartman's Clanker T1000: Local LLM on Framework Desktop with AMD Ryzen AI Max Fuzzing Linux Kernel Bugs
Greg KH's 'gregkh_clanker_t1000' uses a local LLM running on a Framework Desktop (AMD Ryzen AI Max+) to fuzz the Linux kernel, resulting in ~20 merged patches since April 7 fixing bugs in ALSA, HID, SMB, Nouveau, IO_uring, and more.

Claude Security public beta: scans codebase, validates own findings, proposes patches
Anthropic launched Claude Security in public beta for Enterprise customers. It reasons through code like a security researcher, challenges its own findings via adversarial self-verification, and proposes concrete patches.