OpenClaw Agent Broke at 65% Context: 683k Tokens, Zero Cache Reads on Ollama/GLM
A builder ran an OpenClaw agent named Francis in Discord for roughly 23 hours against GLM 5.3 Flash on Ollama Cloud, with a nominal 1,048,576 token context window. About 6 humans across several Discord channels talked to it while it handled tool calls, receipts, coding tasks, code reviews, and file reads. It held together through ~800 transcript events, then failed hard at roughly 65% of the configured context window.
The failure timeline
- ~683k tokens: Francis answered a message about a finished setup task with instructions from a completely unrelated job discussed ~19 hours earlier. Still coherent English, wrong moment.
- 2 minutes later: A long internal planning monologue about the stale task, then collapse into hundreds of repetitions of the word
design. - After that: checkmarks,
tool tool tool,toolResult, fake transcript, repeated numbers, and fragments of its own orchestration envelope. - ~688k tokens: final broken turns.
The session never recovered. Normal controls (/new, /reset, /stop) could not interrupt it, and the operator had to kill the OpenClaw session itself.
The part that actually matters: it reported success
No overflow error. No provider error. No timeout. From the harness's point of view, design design design was a perfectly valid model completion. Automatic compaction was expected around the point where roughly 25% of the window remained — it broke roughly 100k tokens before the safety net was due to fire.
Cache stats: cacheRead=0, cacheWrite=0
Every affected Ollama/GLM turn showed zero visible cache reads and zero cache writes. Cache telemetry was unavailable or zero, so OpenClaw had no evidence the repeated prefix was being reused. Around ten turns in the final 25 minutes each carried a prompt of roughly 680k tokens — about 6.1 million input tokens pushed through in that short window.
A separate agent session shortly before the main collapse recorded 6,912,253 input tokens, 28,956 output tokens, zero cache reads, zero compactions, no timeout, no provider error. Minutes later that agent claimed its conversation history contained fabricated people, tools, and channels, and admitted it could no longer tell which parts of its context were real.
The caching conversation
Minutes before the collapse, the operator asked Francis whether a caching layer was worth building, given all the repeated text moving through the system. Francis confidently said there was nothing useful to do because repeated context is already cheap at the provider via prompt/KV caching. That's a fine answer for OpenAI or Anthropic, where stable prefixes are cached and the telemetry proves it. It was false for the Ollama/GLM route in use — the agent was effectively telling the operator "this bit is cheap" while the harness re-sent ~680k tokens per turn.
The takeaway isn't "small model bad" — the agent did real work for nearly a full day first. The problem is the harness trusted a cache that wasn't there, had no compaction trigger before ~75%, and treated degenerate output as a successful completion. If you're running long-context agents against providers without verifiable cache telemetry, watch your input token accumulation, not just your context percentage.
📖 Read the full source: r/openclaw
👀 See Also

Practical Lessons from Using AI Agents on a 100k LOC Codebase
A developer shares six specific techniques learned while using Claude Code and Cursor to build a pandas-compatible API layer on top of chDB, including maintaining a CLAUDE.md rules file, using zero-context agents as critics, and structuring multi-agent workflows with filesystem-based coordination.

Developer builds LaTeX conversion business in 7 days using Claude Pro
A developer used Claude Pro to build The LaTeX Lab, a service converting Word documents to LaTeX for researchers, in one week for $23.60. The project included market research, AI agent development, custom WordPress theme creation, and SEO-optimized copy.

Practical OpenClaw Setup: Mac Mini Configuration, Cost Management, and Daily Automation
A developer shares their basic OpenClaw assistant setup running on a Mac Mini, detailing security measures, cost optimization from $60-70 initial API fees to $0.60-2.60 daily, and practical integrations including Telegram, Dropbox, and Google Workspace via Composio.

AI Agents Running a Real E-commerce Business: Practical Insights from an Implementation
An AI agent system operates an actual e-commerce store, handling design, coding, marketing, and customer operations without human task execution. The implementation reveals that judgment calls like design rejection thresholds and incident prioritization present harder challenges than technical agent coordination.