Monarch v3: NES-Inspired KV Paging for 78% Faster LLM Inference

What Monarch v3 Does
Monarch v3 is an open-source implementation of NES-inspired memory paging for transformer inference that addresses the linear growth of KV cache with sequence length. By 4K tokens, most KV cache sits unused while consuming VRAM at full precision.
How It Works
The system splits KV cache into two regions:
- Hot region: Recent tokens kept at full precision
- Cold region: Older tokens compressed to ~20 bytes each (vs 64-128 bytes hot)
Four components work together:
- TurboQuant Compression: Quantizes KV to 4-bit integers with polar encoding and residual correction, achieving ~97% size reduction with ~0.3% perplexity loss
- Sliding Window Eviction: Recent N tokens stay hot by default, old tokens compress to cold storage
- Attention-Weighted Promotion: High-attention tokens move back to hot with sticky mechanism to prevent thrashing
- Page Swaps: Small batches of cold tokens materialize on access with local decode loop replacing batch matmul
Benchmark Results
Setup: TinyLlama-1.1B fp16, 50 generated tokens
- Standard: 17.01 tok/s, 2112 MB VRAM
- Monarch-v3: 30.42 tok/s, 2131 MB VRAM, 512 hot tokens, 1024 cold tokens
- Gain: +78.7% throughput, +0.9% VRAM
Simplified Decode Loop
for step in 1..100:
q = project_query(next_token)
# Compute attention: hot only (fast)
scores_hot = q @ kv_hot.T
# Access cold if high attention (rare)
if max(scores_hot) < threshold:
kv_cold_promoted = decompress(cold_pages)
scores_cold = q @ kv_cold_promoted.T
# Move to hot for next step
# Aggregate, softmax, apply attn ...
# Evict old tokens from hot → cold
if len(kv_hot) > window_size:
evict_oldest_to_cold()Current Status
- Implementation: Working on Hugging Face Transformers with custom cache backend
- License: Apache 2.0
- Paper: Full technical spec available
- Next: CUDA kernel fusion for cold decompression planned
How to Try It
git clone https://github.com/JohannaWeb/Monarch.git
cd Monarch
pip install -r requirements.txt
python train_tinyllama_fp16.py
python src/benchmark_monarch.py \
--model models/tinyllama_fp16 \
--mode both \
--max-new-tokens 100 \
--promotion-threshold 0.15 \
--sticky-threshold 3 \
--jsonLimitations
The approach relies on recency (recent tokens = high attention), which works for most tasks but may not for retrieval-heavy workloads. Attention extraction is available in base models but not chat variants; fallback uses window-only paging.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Hollow AgentOSは、JSONネイティブOSアプローチによりClaudeコードトークン使用量を68.5%削減
Hollow AgentOSは、AIエージェント向けに設計されたJSONネイティブのオペレーティングシステム層であり、無駄なシェルコマンドのオーバーヘッドを排除することで、Claude Codeのトークン使用量を68.5%削減します。このツールはMCP経由でClaude Codeに接続し、Ollamaを通じてローカル推論を実行します。

クロードコードがマルチエージェントコードレビューシステムを追加
Anthropicは、プルリクエストをレビューするAIエージェントのチームを派遣するマルチエージェントシステム「Code Review for Claude Code」を発表しました。このシステムは人間のレビュアーが見逃しがちなバグを検出し、実質的なレビューコメントが付くPRの割合が導入前の16%から54%に向上しました。

OpenClawユーザーが、ChatGPTエージェントのワークフロー動作を改善する「feelslikeclaude」スキルを作成しました。
ある開発者がOpenClawのセットアップをClaudeからChatGPTに切り替えたところ、重要な違いは文章のスタイルではなく、ワークフローの振る舞いにあることを発見しました。彼らはChatGPTの実行習慣を改善するために「feelslikeclaude」というclawhubスキルを作成しました。

Cowork Chrome拡張機能がデータブローカーからの個人情報削除を自動化
Redditユーザーによると、Gmailに接続したCowork Chrome拡張機能を使うと、主要データプロバイダーからの個人データ削除リクエストのフォーム入力、メール作成、確認を自動化し、数時間で完了できるそうです。