Developer Tests Qwen3.5 27B vs Larger Models for Local Coding Tasks

✍️ OpenClawRadar📅 Published: March 28, 2026🔗 Source
Developer Tests Qwen3.5 27B vs Larger Models for Local Coding Tasks
Ad

A developer tested several large language models for local coding tasks, comparing performance and hardware requirements. The testing focused on Qwen3.5 variants and Nemotron models, with comparisons to GPT-5.4 High.

Test Results and Findings

The developer tested these specific models:

  • unsloth/Qwen3.5-27B-GGUF:UD-Q4_K_XL
  • unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
  • unsloth/Qwen3.5-122B-A10B-GGUF
  • unsloth/Qwen3.5-27B-GGUF:UD-Q6_K_XL
  • unsloth/Qwen3.5-27B-GGUF:UD-Q8_K_XL
  • unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-IQ4_XS
  • unsloth/gpt-oss-120b-GGUF:F16

Key findings from the testing:

  • Nemotron-3-Super-120B performed "very, very good," on par with GPT-5.4 High
  • Qwen3.5-27B performed well for development tasks
  • GPT-OSS-120B and Qwen3.5-122B performed worse than the other two models
  • Nemotron-3-Super-120B consistently responded in Spanish (the tester's native language) while others responded in English

Performance Metrics

The developer provided specific performance numbers:

  • Nemotron-3-Super-120B: 80 tokens per second (tg/s), ~2000 prompt processing (pp), 100k context on vast.ai with 4x RTX 3090
  • Qwen3.5-27B Q6: 803 pp, 25 tg/s, 256k context on vast.ai
Ad

Hardware Requirements

The developer noted hardware constraints:

  • Qwen3.5-122B would require a new motherboard and 1-2 more RTX 3090 cards, making it too expensive
  • Qwen3.5-27B runs on existing 2x RTX 3090 hardware without additional investment
  • If they had the hardware for Nemotron-3-Super-120B, they would use it instead

Implementation Details

The developer plans to use Qwen3.5-27B-GGUF:UD-Q6_K_XL for real development tasks locally and provided the llama.cpp command used for testing:

./llama.cpp/llama-server -hf unsloth/Qwen3.5-27B-GGUF:UD-Q6_K_XL --ctx-size 262144 --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 -ngl 999

The developer mentioned they'll continue using CODEX for complex tasks but can replace API subscriptions for daily tasks with the local setup.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Storybloq: A Project Tracker for Claude Code with Mac App, CLI, and MCP
Tools

Storybloq: A Project Tracker for Claude Code with Mac App, CLI, and MCP

Storybloq is a free, open-source project tracker that lives in .story/ inside your repo. It includes a Mac app (App Store), a CLI, and an MCP server to expose tickets, issues, and session handovers to Claude Code.

OpenClawRadar
BaseLayer: Open-Source Behavioral Compression Pipeline for AI Memory Systems
Tools

BaseLayer: Open-Source Behavioral Compression Pipeline for AI Memory Systems

BaseLayer is an open-source pipeline that extracts beliefs, behaviors, tensions, and contradictions from conversations, journals, and published text, compressing them into an identity brief for AI models. It has been tested on datasets ranging from 8 personal journal entries to large corpora like Warren Buffett's shareholder letters (350k words) and Howard Marks' investment memos (600k words).

OpenClawRadar
SuperContext: A Persistent Memory Framework for AI Coding Agents
Tools

SuperContext: A Persistent Memory Framework for AI Coding Agents

SuperContext is an open-source framework that gives AI coding tools like Claude persistent memory through structured, targeted files instead of large instruction documents. It includes an executable prompt that builds the system in about 10 minutes with no manual setup.

OpenClawRadar
Codebase Memory MCP: Graph-based code exploration for Claude Code
Tools

Codebase Memory MCP: Graph-based code exploration for Claude Code

A developer built an MCP server that indexes codebases into a persistent knowledge graph using Tree-sitter and SQLite, reducing token usage by 20x on average for structural queries like call tracing and dead code detection.

OpenClawRadar