EvalShift: Open-source CLI for detecting LLM regressions during model migration

EvalShift is an open-source Python CLI designed to detect regressions when switching between LLMs or model versions. It runs your golden input suite against both source and target models, evaluates outputs, and produces a local HTML report — no backend, accounts, or telemetry.
Key features
- Source vs target model comparison via LiteLLM
- JSONL golden suites with tags/slices
- Structural evaluators: JSON schema, regex, length
- Semantic evaluator: embedding similarity
- LLM-as-judge pairwise evaluation
- Tool-call evaluators: tool selection, argument matching, trace structure
- Paired statistical tests: t-test / Wilcoxon
- Effect sizes: Cohen's d
- Multiple-comparison correction: Benjamini-Hochberg
- Slice-level breakdowns
- Local caching to control cost
- Resumable runs
- Single-file HTML report + JSON output
The project's narrow goal is migration safety: “Can I switch models without breaking my prompt/agent behavior?” The author emphasizes catching silent agent regressions — e.g., a newer model producing a decent-looking final answer but skipping a required tool call, calling the wrong tool, or mutating arguments.
Use cases
- Claude 4.5 → Claude 5
- GPT-5 → GPT-6
- Gemini 2 → 3
- Local model → hosted model
The author is seeking feedback on usefulness for local vs hosted models, most important evaluator types for local LLM workflows, and whether tool-call/structured-output regressions are a real pain point. The repo is MIT licensed.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Qwen 3.6 27B F16 Passes Pacman Coding Test, But 8-Bit Quants Fail — Key Lessons on Templates and MTP Speculative Decoding
A user one-shots a Pacman clone with Qwen 3.6 27B F16 — two of three attempts produce nearly perfect games. 8-bit quants fail entirely. Detailed notes on chat template tuning and MTP speculative decoding speed gains.

Claude Code documentation includes excessive React components inflating token counts
Analysis of Claude Code's LLM documentation reveals that MDX files contain massive inlined React components, with context-window.md using 18,501 tokens but only 551 tokens of actual documentation content.

Clawhub Skill Enables OpenClaw to Analyze Apple Health Data via API
A new Clawhub skill called 'apple-health-export-analyzer' allows OpenClaw to read and analyze Apple Health data by serving it as an API, parsing large XML files to extract relevant metrics and provide daily health updates with actionable suggestions.

Claude Code Skill /council Runs Prompts Across 4 AI Models in Parallel
A Claude Code skill called /council sends any prompt to GPT, Claude, Gemini, and Grok simultaneously in about 7 seconds, then uses Gemini to synthesize the best response by identifying specific improvements from the other models.