Auto-Fix System Uses Claude Code Headless to Detect and Fix Production Errors

How the Auto-Fix System Works
A developer has built an automated system that detects and fixes production errors using Claude Code CLI in headless mode. The system has been running for several weeks and is described as free and open source, requiring only a Claude subscription.
System Architecture
The workflow follows this sequence:
- Production logs are monitored
- A watcher fingerprints errors, groups duplicates, and classifies severity
- 30-second settle window
- Critical/High error detection triggers the system
- Git worktree created (isolated branch that never touches main)
- Claude Code launched headless, scoped to the specific error
- Telegram notification: "New Error — Approve Fix?" with Approve/Skip options
- PR created automatically if approved
Key Implementation Details
The developer identified git worktrees as a critical component — each error gets its own isolated copy of the repository. Claude can read, edit, run tests, and perform other operations within this isolated environment. If a fix is unsatisfactory, the worktree can be deleted without affecting the main branch.
Claude sessions receive focused prompts containing:
- Error message
- Stack trace
- Affected path
- Severity level
The headless session runs with scoped tools: Read, Write, Edit, Glob, Grep, and Bash. An example prompt provided: "Fix this production error in the LevProductAdvisor codebase. Error: MongoServerError: connection pool closed. Stack: at MongoClient.connect (mongo-client.ts:88). Path: POST /api/products/list. Severity: CRITICAL."
Results and Performance
According to the developer:
- Critical infrastructure errors (database connection, authentication): Claude fixes 70-80% correctly
- Logic bugs with clear stack traces: Solid performance
- Vague errors without good stack traces: Hit or miss, usually skipped
The system handles straightforward issues like missing null checks or bad query logic effectively, often nailing them on the first attempt.
Additional Features
The developer built an interactive Telegram dashboard for monitoring:
- Queue Status
- Recent Errors
- System Status
- Refresh capability
The /errors view pulls from MongoDB and displays status information like "fixing • 5m ago," "detected • 12m ago," or "fixed • 2h ago."
Technical Stack
The system uses TypeScript, Express, MongoDB, node-telegram-bot-api, and Claude Code CLI. The developer notes that using the headless CLI avoids API costs, requiring only a Claude subscription running locally. Each session is scoped and isolated in a worktree, minimizing risk.
The developer plans to publish the repository on GitHub, describing it as generic — users point the watcher at their log files and configure severity patterns.
📖 Read the full source: r/ClaudeAI
👀 See Also

PACT: A Programmatic Governance Framework for Claude Code After Agent Failure Patterns
A developer built PACT (Programmatic Agent Constraint Toolkit) after three months of recurring Claude Code failures on a 350+ file mobile app. The framework replaces unenforceable rules with mechanical constraints that physically block violations through pre-tool-use hooks.

cc-session-utils: TUI Dashboard for Managing Claude Code Sessions and Costs
A developer built cc-session-utils, a terminal UI tool for managing Claude Code session files, tracking costs by model, cleaning orphaned sessions, and migrating data between projects. It requires Python 3.11+ and is built with Textual.

OpenClaw Skill Reduces Agent Handoff by Enabling Self-Execution
A new skill for OpenClaw agents addresses the common issue where agents identify the next step but stop at 'here's what to do next,' requiring a human handoff. The skill allows agents to carry out certain actions themselves, such as registering, posting, replying, and signing.

agentcache: Python Library for Multi-Agent LLM Prefix Caching
agentcache is a Python library that enables multi-agent LLM frameworks to share cached prompt prefixes, achieving up to 76% cache hit rates and cutting inference time by more than half in tests with GPT-4o-mini.