Token Usage Tips for Claude Code

A detailed Reddit post shares hard-won lessons about managing token consumption in Claude Code. The author notes that most token burn comes from setup context, not Claude's answers. Here are the key practices they use and recommend:
Immediate Wins
- Start new chats for unrelated tasks. Every message in a long conversation resends the full history. A 40-message thread burns tokens on context you stopped caring about 20 messages ago.
- Group small questions into one message. Sending three quick follow-ups individually means three full context loads. Combine them to cut overhead.
- Keep
CLAUDE.mdshort and use it as an index. Dumping everything causes Claude to reread it every turn. Instead, point to separate files so only relevant context loads.
Ongoing Habits
- Be precise with file references. Instead of saying 'here's the whole codebase, figure it out,' which can cost 30–50k tokens in exploration, point Claude to the specific function or module that matters.
- Summarize and restart after 15–20 messages. Ask Claude for a quick summary, paste it into a fresh thread. This drops dead context without losing progress.
- Use lighter models for lighter work. Drafting, reformatting, explaining — route these to smaller models. Reserve the heavy model for reasoning-heavy tasks.
The post invites the community to share their own tricks for keeping token usage under control.
📖 Read the full source: r/ClaudeAI
👀 See Also

Prompt structure improvements for reliable AI skill execution
A developer shares two key prompt modifications that made their market analysis skill run end-to-end without manual intervention: explicitly separating what the skill should return versus what it should do, and defining explicit failure conditions to prevent improvisation.

Using AI to Generate Project Tickets Before Coding Reduces Scope Drift
A developer found that asking AI to generate detailed project tickets with tasks, sub-tasks, scope, and acceptance criteria before writing any code significantly reduced scope creep and large diffs. Each AI agent only receives its specific sub-task, not the entire plan.

llama.cpp Massive Prompt Reprocessing with Coding Agents: Debugging KV Cache and Context Swapping
A user reports llama.cpp reprocessing 40k+ tokens on similar prompts when using opencode + pi.dev, despite high LCP similarity. Config details and suspected causes are shared.
LLM Inference: Techniques for the Efficient Frontier
Basaten's guide to LLM inference engineering: how batch sizing, parallelism, and quantization let you trade latency for throughput or push the entire frontier outward.