Optimizing Qwen3.5-9B on RTX 3070 Mobile with ik_llama.cpp: Config Tweaks and Benchmarks

Hardware and Software Setup
A developer documented their experience optimizing local inference on a laptop with an RTX 3070 Mobile GPU (8GB VRAM, effectively ~7.7GB usable). The system runs CachyOS (Arch-based Linux 6.19) with 32GB RAM and an Intel i7-10750H CPU. They used ik_llama.cpp (ikawrakow's optimized fork of llama.cpp) with the Qwen3.5-9B Q4_K_M model from Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF.
Initial Configuration Issues
The initial naive configuration included several problems:
- MoE-specific flags (
--n-cpu-moe,-ger,-ser) were incorrectly applied to a non-MoE model (n_expert = 0) --mlockwas silently failing due to memory allocation limits (requiresulimit -l unlimitedor limits.conf entry)- Batch size
-b 4096was consuming excessive VRAM (2004 MiB compute buffer), nearly 2GB on an 8GB card
This configuration produced ~47.8 t/s generation speed and ~82 t/s prompt evaluation with VRAM at ~97%.
Optimization Results
After fixing the configuration issues and adjusting batch sizes to -b 2048 -ub 512 (reducing compute buffer to 501 MiB), the developer tested different KV cache configurations:
- Original (q4_0/q4_0, b4096): 47.8 t/s gen, 82.6 t/s prompt, ~97% VRAM
- Fixed flags + b2048/ub512, q8_0K/q4_0V: 48.4 t/s gen, 189.9 t/s prompt, ~80% VRAM
- q8_0K/q8_0V: 50.0 t/s gen, 213.0 t/s prompt, ~84% VRAM
The prompt evaluation speed increased dramatically from ~82 to ~213 t/s, primarily from reducing batch size to free up GPU memory. While generation speed showed minimal change (~2% difference between q4_0 and q8_0), the q8_0/q8_0 configuration produced noticeably more coherent and complete responses on longer outputs, worth the extra ~256 MiB VRAM usage.
Final Configuration
The optimized command for single-user local server use:
./build/bin/llama-server \
-m ./models/Qwen3.5-9B.Q4_K_M.gguf \
-ngl 999 \
-fa on \
-c 65536 \
-b 2048 \
-ub 512 \
-ctk q8_0 \
-ctv q8_0 \
--threads 6 \
--threads-batch 12Open Questions and Future Testing
The developer identified several areas for further investigation:
- GPU power limit tuning on mobile GPUs (potential to reduce TGP with minimal speed loss since inference is memory-bandwidth bound)
- Other 8GB-compatible models with good coding or reasoning performance
- Comparison of ik_llama.cpp vs mainline llama.cpp (ik-specific optimizations include fused ops and graph reuse)
- Tips for hybrid SSM architecture (context shift warnings cause hard stops when context fills, no sliding window)
The testing used a prompt requesting implementation of a Rust Sieve of Eratosthenes program with algorithm explanation, complexity analysis, and example output for N=50.
📖 Read the full source: r/LocalLLaMA
👀 See Also

72단계 클로드 설정 체크리스트: 기본 설정에서 고급 사용자까지
상세한 미디엄 아티클이 Claude 설정을 기본 상태에서 고급 파워 유저 기능으로 전환하는 72단계 체크리스트를 설명합니다. HN에 10점과 1개의 댓글로 공유되었습니다.

클로드와 OpenAI 사용을 위한 모델 라우팅 기준선
한 개발자가 Claude Haiku 4.5, Sonnet 4.6, Opus 4.6 및 ChatGPT 5.3 Codex를 다양한 작업 유형에 사용하는 모델 라우팅 전략을 공유하며, 필요 시 GPT-5 Mini와 GPT-5.4로 폴백합니다.

연구에 따르면 효과적인 AI 프롬프트 작성은 공학적 접근이 아닌 협력적 소통이다
동료 검토 연구에 따르면 AI 모델과의 효과적인 프롬프팅은 인간이 사용하는 협력적 의사소통 원칙과 동일하게 작동하며, Lakera의 분석에 따르면 대부분의 프롬프트 실패는 모델의 한계가 아닌 모호함에서 비롯됩니다.

150개 이상의 PR/주로 에이전트 코딩 확장: Lovable에서 토큰 85,000달러 사용으로 얻은 교훈
Alexander Lebedev가 2026년 1월부터 AI 토큰에 85,000달러를 사용하여 주당 20~30개의 PR에서 150개 이상의 PR로 확장한 방법을 공유합니다. 주요 학습 포인트: 위험 분류, AI 리뷰가 인간 코드 리뷰를 대체, 지식 확산 보존의 과제.