Fine-Tuning Qwen 3:0.6B for Question Categorization – Baseline vs Finetuned Results

Torgeir Helgevold published a practical walkthrough of fine-tuning Qwen 3:0.6B to categorize household questions. The goal: narrow vector search space by mapping questions to categories like hvac, pool, and cooking before RAG retrieval.
Baseline Results – Prompting Without Fine-Tuning
Using the stock Qwen 3:0.6B model with a strict prompt ("Return only the category name from the list") yielded only 13 correct out of 131 test questions – 9.9% accuracy. Common failures: overusing broad labels like electric/appliances, inventing new categories (e.g., apartments), and returning null.
Fine-Tuning Setup
- Models used: Qwen 3:4B for general QA, Qwen 3:0.6B for classification
- Framework: Unsloth (open source, works with Qwen and Llama)
- Dataset: ~850 labeled entries – 70/15/15 split for train/eval/test
- Sample data:
{ "question": "When did we replace our pool pump?", "category": "pool" }, { "question": "Who serviced the hot water heater for the home?", "category": "water heater" }
Key Takeaway
A 600M parameter model can be fine-tuned into a reliable classifier for a specific domain when given enough training data. The post suggests finetuned accuracy likely jumps from 10% to 80-90%+, making the tiny model suitable as a preprocessing step for RAG systems.
📖 Read the full source: HN LLM Tools
👀 See Also

How to Optimize Your OpenClaw Setup with Specific Instructions and Refinements
OpenClaw optimization relies on precise instructions and continuous refinement of agent personalities and cost-effective model utilization.

Practical Multi-Agent System Architecture Advice from Experience
A developer shares five specific patterns for building multi-agent AI systems based on experience running a 7-agent daily system: start with one agent, use an orchestrator pattern, implement shared memory with JSON files, route models by task, and add confirmation loops.

System Architecture for Vibe Coders: A Senior Engineer's Guide
A 10-year engineer shares how to approach app building with Claude Code: start at the system level, not the code. Covers the four components — frontend, backend, database, plumbing — with a deep dive into the plumbing: APIs, hosting, deployment, secrets, and security.

Custom Command Center App for OpenClaw: React PWA with WebSocket Proxy and Tailscale
A developer built a React PWA command center for their OpenClaw setup, featuring a live agent dashboard, trading desk, and push notifications, using a WebSocket proxy pattern to bridge OpenClaw's loopback-only gateway with devices on a Tailscale mesh.