Open-source local hook automatically switches Claude models to cut AI costs

A developer has open-sourced a local hook that automatically selects the most cost-effective Claude AI model based on the type of coding task, potentially reducing AI costs by 50-70% without quality loss.
How it works
The tool runs as a local hook in Cursor and Claude Code (both use the same hook system) before each prompt is sent. It sits next to Opus/plan and acts as an efficient front-end filter that prevents obviously bad model matches before they hit expensive models.
Key functionality
- Reads the prompt and current model selection
- Uses simple keyword rules to classify tasks (git operations, feature work, architecture/deep analysis)
- Blocks if you're overpaying (e.g., Opus for git commit) and suggests Haiku or Sonnet
- Blocks if you're underpowered (Sonnet/Haiku for architecture) and suggests Opus
- Lets everything else through unchanged
- ! prefix bypasses the filter completely if you disagree with its suggestion
Technical details
- 3 files: bash + python3 + JSON
- No proxy, no API calls, no external services
- Fail-open design: if it hangs, Claude Code proceeds normally
- Open-sourced at: https://github.com/coyvalyss1/model-matchmaker
Performance and testing
The developer analyzed several weeks of their own prompts and found:
- 60-70% were standard feature work Sonnet could handle
- 5-20% were debugging/troubleshooting
- A significant portion were pure git/rename/formatting tasks that Haiku handles identically at 90% less cost
Retroactive analysis showed the tool would have cut 50-70% of AI spend with no quality drop. After tuning, it correctly handled 12/12 real test prompts.
Problem it solves
The issue isn't knowledge—developers know they should switch models—but friction. When in flow state, developers don't want to think about dropdown menus. This tool automates the decision-making process.
📖 Read the full source: r/ClaudeAI
👀 See Also
Researcher Builds Veracity-Checking Skill for Claude Code, Finds Hallucinations in Own Documentation
A researcher built a Claude Code skill called /veracity-tweaked-555 that decomposes documents into atomic claims and verifies each via web search using 16 parallel agents across 4 waves. When self-audited, the skill scored 62/100 due to fabricated statistics and inflated claims in its own documentation.

Claude wrote 3,000 lines of code instead of importing pywikibot — a case study in AI agents ignoring existing libraries
A developer tasked Claude Code (Opus 4.7) with fixing typos on Fandom wikis. The model wrote ~3,000 lines of Python reimplementing pywikibot, mwparserfromhell, and RETF rules rather than importing them. The post explores why this happens and how a two-minute search reduced the codebase to 1,259 lines.

Jeeves: TUI for Browsing and Resuming AI Agent Sessions
Jeeves is a terminal user interface that lets you search, preview, and resume AI agent sessions from Claude Code, Codex, and OpenCode in a single view. It's written in Go and available via multiple package managers including Homebrew, Nix, and Go install.

Developer Creates Practical Claude Skills for Kotlin Multiplatform Projects
A developer built a public repository of Claude skills specifically for Kotlin Multiplatform work after finding existing skills too generic, opinionated, or thin. The skills cover architecture reviews, feature implementation, modularization, Compose Multiplatform UI, navigation, platform bridges, deep links, adaptive UI, testing, and build governance.