New AI Tutor Achieves 0.71-1.30 SD Effect Size in Dartmouth Course

Researchers at Dartmouth deployed an AI tutor in an introductory computer science course and measured effect sizes of 0.71 to 1.30 standard deviations on learning outcomes. The paper, presented at the 2026 InTextbooks workshop, compares the AI tutor to standard instruction. The results suggest that LLM-based tutoring can substantially outperform traditional methods in controlled classroom settings.
Key Findings
- Effect size range: 0.71 to 1.30 SD across different assessment types
- Study conducted in a Dartmouth introductory CS course
- AI tutor likely leverages Socratic-style hints and code feedback via LLM
- Control group received standard instruction without the AI tutor
While the exact architecture is not fully detailed in the PDF snippet, the effect sizes are large enough to be practically significant. An effect size of 1.0 SD typically corresponds to moving an average student from the 50th to about the 84th percentile.
Who This Matters For
Developers building educational agents or tutoring systems for coding. Also relevant for AI researchers evaluating real-world LLM impact.
📖 Read the full source: HN LLM Tools
👀 See Also

Investigation: Claude Code Agents Surfacing Unverified MEMORY.md Content Due to Compaction Changes
A user reports that Claude Code agents are surfacing content from MEMORY.md without re-verifying mid-task, linked to compaction changes in versions 2.1.139 and 2.1.141. Two compounding factors: aggressive preservation of 'user instructions' and a bug in autocompact thresholds.
Claude Code v2.1.252 Fixes Bash Errors, Remote Control Stalls
Claude Code v2.1.252 fixes Bash failures on Macs, 'always allow' persistence, Remote Control stalls in Claude Desktop/VS Code, and oversized background notifications.
NVIDIA Groq 3 LPX Hits 3,431 Tokens/sec on Long Context Benchmarks
NVIDIA's Groq 3 LPX accelerator, paired with Vera Rubin NVL72, achieves 3,431 output tokens per second on the 100K context benchmark with Gemma 4 31B, enabling high-interactivity inference for multi-turn agentic workloads.

Anthropic blocks third-party harnesses from Claude subscription limits, workaround available
Anthropic has restricted third-party harnesses from accessing Claude subscription limits, potentially disrupting workflows that rely on these tools. A Reddit user reports developing an open-source workaround after nearly losing months of training data.