agentcache: Python Library for Multi-Agent LLM Prefix Caching

agentcache is a Python library designed to optimize multi-agent LLM systems by implementing prefix caching as a core feature. The library addresses the common problem where frameworks like CrewAI, AutoGen, and open-multi-agent create fresh sessions for each worker, resulting in zero cache hits and duplicated prompt costs.
How It Works
The library operates on a fork-based approach instead of creating separate sessions:
- Start one session with a shared system prompt
- Make the first call - provider computes and caches the prefix
- When you need N workers, fork instead of creating N new sessions
- Parent session: [system, msg1, msg2, ...]
- Forked session: [system, msg1, msg2, ..., WORKER_TASK]
- Exact same prefix = cache hit
Key Features
- Cache-safe forks: Maintains identical prefixes across worker sessions
- Cache-break detection: Diffs snapshots and reports exactly what changed when cache hits drop
- Cache-safe compaction: For long-running sessions, scans old tool outputs before each call and replaces large results with deterministic placeholders to maintain smaller context while preserving cacheable prefixes
- Parameter freezing: Freezes cache-relevant parameters before forking (system prompt, model, tools, messages, reasoning config)
- Task DAG scheduling: Enables parallel workers from one cached session
Performance Results
In a head-to-head test with GPT-4o-mini (coordinator + 3 workers, same task):
- Text injection / separate sessions: 0% cache hits, 85.7 seconds
- Prefix forks: 75.8% cache hits, 37.4 seconds
- Per worker cache hit rates typically range from 80-99%
Installation and Usage
Install via pip:
pip install "git+https://github.com/masteragentcoder/agentcache.git@main"
The library is available on GitHub at github.com/masteragentcoder/agentcache.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Localhost dashboard for Claude's HTML output pages with live reload
A small localhost home page collects all HTML pages generated by Claude AI, with live reload, update indicators, and a preview pane. Agents register pages via symlinks.

Two Free Claude Code Skills: Tutorial Generator and Prompt Fixer
Two new free Claude Code skills: create-tutorial generates code reading tutorials from your actual project files, and prompter rewrites typo-filled prompts into actionable instructions. Both are MIT licensed and install via GitHub.

TasteBud Memory: Reversible Agent Memory via Hyperdimensional Computing
A 600-line Node.js tool uses hyperdimensional computing to build a reversible memory layer for AI agents, supporting lossless decode, drift detection, and unknown-project alerts.

GitHub Comic Bot: Turn Commits into Daily Medieval Knight Comics
A bot that reads GitHub commits and generates 4-panel comic strips featuring a deadpan medieval knight, built with Claude Code and Gemini, running on GitHub Actions with free tier costs.