When RLVR Helps Small Fine-Tuned Models: A 12-Dataset Analysis

A recent experiment tested whether adding a reinforcement learning stage (RLVR) on top of supervised fine-tuning (SFT) for small language models (1.7B parameters) provides measurable benefits. The team ran a controlled experiment across 12 datasets to determine exactly when this approach helps and when it doesn't.
Key Findings
The results split cleanly by task type:
- Text generation tasks (QA, documentation, PII redaction): +2.0 percentage points average improvement. Every single dataset in this category showed improvement.
- Structured tasks (classification, function calling): -0.7 percentage points average. Two datasets in this category actually regressed.
Why This Pattern Emerges
The researchers explain that once a fine-tuned model already gets most structured outputs correct, GRPO (Group Relative Policy Optimization) produces near-zero gradients. Essentially, there's no learning signal left for the reinforcement learning stage to work with.
For generative tasks, the output space is large enough that RL continues to find improvements that SFT misses — particularly when rewarding semantic correctness rather than exact string matching.
Practical Decision Rule
The study provides a simple guideline for developers:
- Classification or strict function calling → Use SFT only
- QA, documentation, extraction tasks → Add RLVR on top of SFT
The methodology, all 12 datasets tested, and raw numbers are available in the full analysis.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Inference Pricing Analysis Shows 4.4x Spread for Same Model Across Providers
Analysis of inference pricing for Llama 3.1 70B Instruct shows a 4.4x cost difference between providers, with DeepInfra at $0.20/$0.27 per million tokens and Together at $0.88/$0.88. For reasoning models, the spread reaches ~30x between DeepSeek R1 and OpenAI o1.

OpenClaw v2026.7.1: Control UI Overhaul, Onboarding, Mobile Apps, GPT-5.6, Tencent Hy3, Meta Muse Spark 1.1
OpenClaw v2026.7.1 brings a major Control UI overhaul, redesigned onboarding, updated iOS/Android/macOS apps, GPT-5.6 compatibility, Tencent Hy3 and Meta Muse Spark 1.1 support, and improved Codex and coding-agent workflows.

China Bars Manus Co-Founders from Leaving Country Amid Meta Deal Review
China has barred two co-founders of AI startup Manus from leaving the country as regulators review whether Meta's $2 billion acquisition violated investment rules. The executives were summoned to Beijing for a meeting with the National Development and Reform Commission this month.

Claude Code v2.1.37 Released
Anthropic releases a new version of Claude Code with improvements and bug fixes.