Qwen3.6 27B FP8 Runs 200k Tokens BF16 KV Cache at 80 TPS on RTX 5000 PRO 48GB

A Reddit user on r/LocalLLaMA reports running Qwen3.6-27B-FP8 with a BF16 KV cache of 200k tokens at 60–90 TPS on a single RTX 5000 PRO 48GB GPU. The setup uses vLLM 0.20.1, CUDA 12.9, and Qwen's official FP8 quant, preserving multi-modality and MTP speculative decoding.
Setup Details
The environment uses FlashInfer FP8 MoE, FP8 Marlin, and async scheduling. Key environment variables and launch command:
export VLLM_USE_FLASHINFER_MOE_FP8=1
export VLLM_TEST_FORCE_FP8_MARLIN=1
export VLLM_SLEEP_WHEN_IDLE=1
export VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=1
export VLLM_LOG_STATS_INTERVAL=2
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export SAFETENSORS_FAST_GPU=1
export CUDA_DEVICE_ORDER=PCI_BUS_ID
export TORCH_FLOAT32_MATMUL_PRECISION=high
export PYTORCH_ALLOC_CONF=expandable_segments:True
vllm serve Qwen/Qwen3.6-27B-FP8
--host 0.0.0.0 --port 8080
--performance-mode interactivity
--trust-remote-code
--enable-auto-tool-choice
--tool-call-parser qwen3_coder
--reasoning-parser qwen3
--mm-encoder-tp-mode data
--mm-processor-cache-type shm
--gpu-memory-utilization 0.975
--speculative-config '{"method":"mtp","num_speculative_tokens":2}'
--compilation-config '{"cudagraph_mode": "FULL_AND_PIECEWISE", "max_cudagraph_capture_size": 16, "mode": "VLLM_COMPILE"}'
--async-scheduling
--attention-backend flashinfer
--max-model-len 196608
--kv-cache-dtype bfloat16
--enable-prefix-caching
Performance Observations
With MTP=2 speculative decoding, the system produces 60–90 TPS during code generation. The BF16 KV cache avoids compaction issues seen in quantized KV, making long coding sessions more reliable. The user notes that the setup runs on a single RTX 5000 PRO 48GB with a 64GB system RAM and a decent CPU, calling it a strong candidate for a $10k workstation for local LLM development.
Who It's For
Developers needing a local, low-compression agentic coding setup with minimal quantization artifacts and long context windows.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Claude Code v2.1.225 Fixes OAuth Token Rotation, Adds Gateway Spend-Limit Support
v2.1.225 fixes a transient 401 that broke headless sessions, adds gateway spend-limit support, and improves Remote Control features.

The Frontier AI Race is Over: Networks of Smaller Models Beat Centralized AI on Cost and Capability
Networks of smaller AI models now outperform every frontier AI system on speed, accuracy, and cost. The article argues that centralized AI companies cannot regain the lead due to the "Hydra Effect" — ensembling cheaper models recursively beats any single model.

CEOs Who Think AI Replaces Their Employees Are Just Bad CEOs
CEO Aaron Levie explains 'AI psychosis' — when leaders, detached from real work, see happy-path demos and overestimate agentic tools like Claude Code, ignoring the last mile of production.
Claude Code v2.1.275: Send-Now Key, claude.ai Skill Sync, and Two Dozen Bug Fixes
Claude Code v2.1.275 adds a ctrl+enter send-now key that flushes queued messages, syncs claude.ai skills and plugins to terminal sessions, and fixes prompt cache misses plus malformed-transcript resume crashes.