llama.cpp Massive Prompt Reprocessing with Coding Agents: Debugging KV Cache and Context Swapping

A developer on r/LocalLLaMA is hitting a serious performance issue with llama.cpp when running long-context coding agents (opencode + pi.dev) via llama-swap. Even with highly similar prompts (LCP similarity often >0.99), the system periodically discards the KV cache and reprocesses 40k+ tokens, causing TTFT of multiple minutes.
Observed Behavior
- Context grows to 50k+ tokens.
- After several normal reuses (e.g.,
prompt eval time = 473 ms / 19 tokens),n_pastsuddenly drops to ~4-5k. - llama.cpp then reprocesses the full prompt:
n_tokens = 4750 prompt eval time = 222411 ms / 44016 tokens. - Cache usage hits 4676 MiB, exceeding the configured limit (2500 MiB).
Current Configuration
llama-server --ctx-size 150000 --parallel 1 --ctx-checkpoints 32 --cache-ram 2500 --cache-reuse 256 -no-kvu --no-context-shiftSuspected Causes
- Cache invalidation due to overflow of
--cache-ramlimit – the log shows 4676 MiB used vs 2500 MiB limit. - Bad KV reuse mechanism when early prompt tokens change (possibly frequent alterations by opencode).
- Insufficient
--ctx-checkpointsor--cache-reusefor the 150k context size.
Recommendations from the Community
The thread is thin on answers so far, but obvious first steps include increasing --cache-ram to match typical usage (e.g., 5000+ MiB), or reducing --ctx-size to stay under the cache limit. Also check if opencode is intentionally mutating prompt prefixes; if so, locking the system prompt or using a fixed prefix could improve reuse.
For developers running similar setups, share your working configs in the source thread.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Workaround for Control UI assets error after OpenClaw 2026.3.22 upgrade
A user posted a solution for the 'Control UI assets not found' error that occurs after upgrading to OpenClaw 2026.3.22, involving copying the control-ui folder from a beta installation to the stable release.

Vague Prompts Are the Real Problem, Not the Model — 50-Run Test Shows Prompt Quality Trumps Model Choice
A Reddit user ran the same ten prompts through ChatGPT 4, Claude Sonnet, and Gemini 1.5 Pro five times each (150 outputs total) and found that all three models produced similarly usable or similarly generic results — the deciding factor was prompt specificity, not the model.

Prompt structure improvements for reliable AI skill execution
A developer shares two key prompt modifications that made their market analysis skill run end-to-end without manual intervention: explicitly separating what the skill should return versus what it should do, and defining explicit failure conditions to prevent improvisation.

Stop Copy-Pasting Errors Into Claude Code — Give It Access Instead
Don't copy-paste errors into Claude Code. Instead, give it the API keys or tools it needs to self-diagnose and fix. The author shares practical patterns for staging databases, headless browsers, and eval environments.