Fine-tuned Qwen2.5-7B to 96% of Claude Haiku with $3 and Zero Human Labelers

A developer fine-tuned Qwen2.5-7B to achieve 96% of Claude Haiku's composite performance on a domain-specific decision-reasoning task — spending only ~$3 in API calls and using zero human labelers. The method, called DV-DPO (Decision-Validated Direct Preference Optimization), autonomously generates training signal by running a multi-voice adversarial council.
How DV-DPO Works
The pipeline runs a 3-voice council on each decision question, producing a synthesis. Then the two losing voices cross-examine the synthesis. If the synthesis is revised under this adversarial pressure, a DPO pair is formed: the post-revision version is the chosen response, and the pre-revision version is the rejected response. If the synthesis holds — no pair is created. This ensures only genuine reasoning errors produce training signal, not format preferences or sampling variance.
Results
- 1,040 training pairs generated total (~$3 at Haiku rates)
- Head-to-head vs Claude Haiku: Format 100%, Commits 100%, Context 89%, Composite 96%
- Latency: 11s on T4 GPU (4-bit quantized) vs Haiku's 3s
- Adversarial failure rate: 2% on 96 targeted questions
Autonomous Improvement Loop
The system now runs an automated cycle: failure_detector → auto_red_team → DPO pairs → retrain → redeploy → eval. Version 5 pairs are accumulating. The fine-tuned model is available as a GGUF file ready for Ollama.
Who This Is For
Developers building domain-specific reasoning agents who want to move from pay-per-call APIs to a local fine-tuned model without expensive human annotation.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Open-source models match or beat Claude Opus 4.6 on benchmarks
DeepSeek V3.2, DeepSeek R1, Kimi K2.5, and MiniMax M2.5 outperform Claude Opus 4.6 on 4 out of 5 major benchmarks including MMLU-Pro, speed, tool use, and reasoning, while being significantly cheaper.

In AI, the 41% Depends on the -59%
Apollo's chief economist breaks down AI profit margins by value chain layer: Models & Apps run at -59% while upstream Silicon & Equipment hit 41%. The boom is funded by investors, not customers.

Claude Code v2.1.74 System Prompt Updates: Security Rules, Memory Selection, and New Skills
Claude Code v2.1.74 adds 1,750 tokens to system prompts including new security monitor rules blocking unauthorized external writes, a /stuck skill for diagnosing frozen sessions, and memory selection improvements that skip redundant API references.

CEOs Who Think AI Replaces Their Employees Are Just Bad CEOs
CEO Aaron Levie explains 'AI psychosis' — when leaders, detached from real work, see happy-path demos and overestimate agentic tools like Claude Code, ignoring the last mile of production.