Optimizing Qwen 3.6 27B/35B on RTX 3090: Flags, Quantization, and Auto-Routing

A developer running Qwen 3.6 models locally on an RTX 3090 (24GB VRAM), Ryzen 5700X, 64GB RAM, Windows 11, is hitting performance and reliability issues. They're using llama-server with custom flags and seeking advice on quant choice, throughput, and automatic model routing.
Commands and Quantizations
35B (UD Q4_K_M):
llama-server.exe -m "path\Qwen3.6-35B-A3B-UD-Q4_K_M.gguf" -ngl 99 -c 131072 -np 2 -fa on -ctk f16 -ctv f16 -b 2048 -ub 512 -t 8 --mlock -rea on --reasoning-budget 2048 --reasoning-format deepseek --jinja --metrics --slots --port 8081 --host 0.0.0.027B (UD Q4_K_XL):
llama-server.exe -m "path\Qwen3.6-27B-UD-Q4_K_XL.gguf" -ngl 99 -c 196608 -np 1 -fa on -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 8 --no-mmap -rea on --reasoning-budget -1 --reasoning-format deepseek --jinja --metrics --slots --port 8081 --host 0.0.0.0Reported Issues
- 35B too slow – even simple iterative tasks feel unusable.
- 27B faster but unreliable – code output breaks; simple tasks can take 20–30 minutes.
- Manual model switching – must kill server, paste new command, reload model.
Specific Questions
- Are the flags suboptimal? (e.g., context size, batch size, cache type)
- Which quant / model gives best balance of speed and coding accuracy on 24GB VRAM?
- How to auto-switch models per request, or keep multiple models warm and route?
Context
The user runs Hermes agent on a Raspberry Pi 5 for scraping and automation, and local coding with OpenCode/QwenCode. They want a setup that doesn't require manual server restarts.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Fix for Claude Desktop Workspace VM Service Issue on Windows 11 Home
A community-developed fix addresses the 'VM service not running' error in Claude Desktop's workspace feature on Windows 11 Home, with manual PowerShell commands and an automated tool available on GitHub.

Running Qwen3.6-35B-A3B with ~190k Context on 8GB VRAM + 32GB RAM – Setup & Benchmarks
A Reddit user shares a working llama.cpp configuration for Qwen3.6-35B-A3B GGUF models on an RTX 4060 (8GB VRAM) + 32GB DDR5, achieving 37-51 tok/s at 192k context using TurboQuant and specific flags.

Skill-writing principles for Claude Code from 159 open-source skills
A developer shares 10 principles for writing effective skills for Claude Code, based on building and maintaining an open-source registry with 159 skills. The principles include practical approaches like using folders instead of single files, adding gotchas sections, and implementing on-demand hooks.

Getting the Most Out of Claude: A Data Analyst's Workflow with Cowork and Claude Code
A data analyst with no coding background shares how they use Cowork for end-to-end automation and Claude Code for heavy lifting — building a lead gen tool using Google Places API, a fraud dashboard, and automated social media posting.