Claude AI Session Compaction Issues and Workarounds

How Compaction Works
Claude sessions are stored as JSONL files at ~/.claude/projects/{encoded-cwd}/sessions/{id}.jsonl. Each conversation turn is a JSON block. When compaction triggers, original blocks remain in the file, but a new block with a compressed summary gets appended. After compaction, the model works from the summary instead of the full conversation history.
Test Results
With a coding project at 90% context fill (before the 1 million token increase), the user tested 10 questions covering simple recall, 6-hop dependency chains, entity disambiguation, negation chaining, absence detection, and conflict detection.
- Pre-compaction: ~9.75/10 accuracy with Opus 4.6 finding scattered facts across 418K tokens
- Post-compaction (Default): ~5/10 accuracy with 3,461 tokens (121x compression). Same session, same questions resulted in hallucinated incorrect answers.
- Post-compaction (Manual Opus): ~9.75/10 accuracy with 6,080 tokens (69x compression). Using a custom compaction prompt with Opus preserved important information.
Why the Difference
According to Anthropic's documentation, the API defaults to using the same model for compaction. The user was running Opus 4.6 on medium compute, so default compaction should have used Opus too. The quality difference suggests issues with the summarization prompt, thinking/compute budget, or both.
Workarounds
Approach 1: Opus Compaction - Turn off auto-compaction and implement a background process that measures token counts for Claude Code instances. Trigger compaction using Opus with a custom prompt (potentially with user authorization).
Approach 2: spaCy NER Pre-seeding - Instead of starting sub-agents with zero context, use spaCy NER to extract proper nouns, numbers, service names, ports, and key identifiers from project files. Inject this as a lightweight entity briefing (few hundred tokens) at startup to inform agents about existing resources without narrative bloat.
📖 Read the full source: r/ClaudeAI
👀 See Also

Persistent Memory for Claude: Local Stack with MCP, 39ms Retrieval, 82% Token Reduction
A developer built a persistent memory layer for Claude using local vector search (Qdrant + Qwen3) and MCP integration, achieving 82% token reduction, 39ms hot-path retrieval, and session crystallization via L4 nodes.

memora: Version-Controlled, Typed Memory for AI Agents – Git for AI Beliefs
memora is a CLI tool written in Rust that version-controls AI agent memory with typed, provenance-tracked, branchable, and mergeable capabilities.

Claude-First Analytics MCP Server: Giving AI Agents Direct Access to Web Analytics Context
A developer rebuilt their web analytics tool as an MCP server, exposing simple web analytics, trackable links, and product insight tools directly to Claude, enabling AI agents to leverage site data alongside code and database context.

Publicly Hosted MCP Servers for Health, Academic, and Government Data
A developer has built and publicly hosts 14 MCP servers providing access to CDC datasets, clinical trials, FDA data, academic publications, congressional information, weather data, and other utilities. These servers require no setup, API keys, or local installation.