Pi Coding Agent with Qwen 35B Q2: Using Filesystem as External Memory and Enforcing Context Guards

A Reddit user shared their approach to agentic coding with local LLMs, built on Pi coding agent with Qwen 35B (Q2_K_XL quant via LM Studio). The core insight: treat the LLM as a logic processor, not a context database. The implementation enforces strict guards at the API boundary — the model cannot bypass them.
Key constraints enforced by the system
- Write/edit limit: Rejects any output over 100 lines. Model must write a skeleton first, then fill in one section at a time. If it tries to dump a full file, the call is blocked with instructions to split the work.
- Thinking block cap: If the model's reasoning exceeds 2000 chars, it receives a correction to write conclusions to disk and move on.
- Context monitor: At 65% context usage, the model is told to write its state to files. At 80%, everything stops — the model writes its 'brain' to disk while still coherent.
- Persistent output: If the model gives a long answer without writing a file, it's instructed to save findings to a step file. Nothing stays only in context.
External brain structure
The system uses .think/ and .plan/ directories as the model's external memory. Every step, decision, and finding is written to a file. When context compresses, the model reads its own notes back. The session purpose is saved separately to _purpose.md and re-injected after context compression, preserving the original goal.
Knowledge distillation
A /distill command crawls a codebase, builds an import graph, topologically sorts files, and has the model summarize them one per turn into a knowledge base. The manifest is split into pages of 50 files to avoid consuming the whole context. Users can drop files like svelte5-gotchas.md or astro-gotchas.md into a knowledge folder; an isolated LLM call selects which ones are relevant to the current task, and only the content gets injected into the main conversation.
Real-world result
The user asked the model to build a Three.js plane flying game. The first attempt tried to write 652 lines in one call — the guard rejected it. The model replanned, wrote a skeleton, then filled in features one edit at a time. The final result was a working game with 3D plane model, obstacles, HUD, minimap, and start/game over screens — all at Q2 quant.
The full setup runs at Q2_K_XL quantization as the floor; the user notes Q4 or Q8 should yield better results. The code is available on GitHub: github.com/Kodrack/Pi-forge.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Doc Harness: A Claude Code Skill for Maintaining Project State Across Sessions
Doc Harness is a Claude Code skill that creates a lightweight documentation system with five structured files to help AI agents maintain project context across sessions. It addresses issues like context resets, forgotten rules, and the need to re-explain projects to new agents.

Lightfeed Extractor: TypeScript Library for Robust Web Data Extraction with LLMs
Lightfeed Extractor is a TypeScript library that handles the full pipeline from raw HTML to validated structured data using LLMs, with features like HTML-to-markdown conversion, Zod schema validation, JSON recovery, and built-in Playwright browser automation.
Cue AI Uses Gemma 4 for Faster Voice Dictation: 44% Latency Drop, 30% More Usage
Cue AI replaced a cloud-based text polish step with Google DeepMind's Gemma 4 E4B running locally via Ollama, cutting median latency from 876ms to 488ms and increasing dictation usage by 30%.
Ant Group's AntLing-3.0-flash Hits OpenClaw via OpenRouter – 256K Context, Free Through Aug 3
AntLing-3.0-flash from Ant Group is now available on OpenClaw via OpenRouter. No client update needed; configure with openclaw models set openrouter/inclusionai/ling-3.0-flash. Features 256K context and RL training on long-horizon tool calling. Free through Aug 3 (API-only, no weights).