Monarch v3: NES-Inspired KV Paging for 78% Faster LLM Inference

What Monarch v3 Does
Monarch v3 is an open-source implementation of NES-inspired memory paging for transformer inference that addresses the linear growth of KV cache with sequence length. By 4K tokens, most KV cache sits unused while consuming VRAM at full precision.
How It Works
The system splits KV cache into two regions:
- Hot region: Recent tokens kept at full precision
- Cold region: Older tokens compressed to ~20 bytes each (vs 64-128 bytes hot)
Four components work together:
- TurboQuant Compression: Quantizes KV to 4-bit integers with polar encoding and residual correction, achieving ~97% size reduction with ~0.3% perplexity loss
- Sliding Window Eviction: Recent N tokens stay hot by default, old tokens compress to cold storage
- Attention-Weighted Promotion: High-attention tokens move back to hot with sticky mechanism to prevent thrashing
- Page Swaps: Small batches of cold tokens materialize on access with local decode loop replacing batch matmul
Benchmark Results
Setup: TinyLlama-1.1B fp16, 50 generated tokens
- Standard: 17.01 tok/s, 2112 MB VRAM
- Monarch-v3: 30.42 tok/s, 2131 MB VRAM, 512 hot tokens, 1024 cold tokens
- Gain: +78.7% throughput, +0.9% VRAM
Simplified Decode Loop
for step in 1..100:
q = project_query(next_token)
# Compute attention: hot only (fast)
scores_hot = q @ kv_hot.T
# Access cold if high attention (rare)
if max(scores_hot) < threshold:
kv_cold_promoted = decompress(cold_pages)
scores_cold = q @ kv_cold_promoted.T
# Move to hot for next step
# Aggregate, softmax, apply attn ...
# Evict old tokens from hot → cold
if len(kv_hot) > window_size:
evict_oldest_to_cold()Current Status
- Implementation: Working on Hugging Face Transformers with custom cache backend
- License: Apache 2.0
- Paper: Full technical spec available
- Next: CUDA kernel fusion for cold decompression planned
How to Try It
git clone https://github.com/JohannaWeb/Monarch.git
cd Monarch
pip install -r requirements.txt
python train_tinyllama_fp16.py
python src/benchmark_monarch.py \
--model models/tinyllama_fp16 \
--mode both \
--max-new-tokens 100 \
--promotion-threshold 0.15 \
--sticky-threshold 3 \
--jsonLimitations
The approach relies on recency (recent tokens = high attention), which works for most tasks but may not for retrieval-heavy workloads. Attention extraction is available in base models but not chat variants; fallback uses window-only paging.
📖 Read the full source: r/LocalLLaMA
👀 See Also

5090에서 Qwen3.6-27B와 Opencode를 이용한 로컬 AI 개발
한 Reddit 사용자가 클라우드 AI 코딩 도구(Claude Code, Cursor)에서 로컬 설정(Opencode + llama-server + Qwen3.6-27B, 128K 컨텍스트, 단일 RTX 5090)으로 전환한 경험을 공유하며, 사용량 제한과 계정 위험에서 자유로워졌다고 말합니다.

자율적 클로드 코드 세션을 위한 디스코드 브리지
약 50줄의 bridge.js 스크립트(discord.js v14)가 WebSocket과 로컬 파일 큐를 통해 Discord와 Claude Code 간 실시간 양방향 채팅을 생성하며, 2분 간격 폴링을 마이크로초 단위 파일 읽기로 대체합니다. 27,000줄을 밤새 분석하여 테스트되었습니다.

카패시 코딩 스킬, 프리 플랜으로 재탄생…프로 없이 클로드 코딩 훈련 해제
Reddit 사용자가 Karpathy의 코딩 규율 가이드라인을 Claude의 무료 플랜용으로 다시 작성하여 터미널 및 하위 에이전트 의존성을 제거했습니다. 시스템 프롬프트는 코딩 요청 시 자동으로 트리거되고 검증 우선 사고를 강제합니다.

Forge: Mac 또는 Linux 머신을 AI 코딩 에이전트를 위한 항상 켜진 개발 호스트로 전환하세요
Forge는 데몬을 설치하여 Mac 또는 Linux 머신을 영구적이고 항상 켜진 개발 호스트로 전환하는 오픈소스 도구입니다. 사용자가 자리를 비울 때 AI 코딩 에이전트를 계속 실행하고, 모니터링을 위한 웹 대시보드를 제공하며, Tailscale을 사용하여 SSH를 통한 안전한 원격 접속을 가능하게 합니다.