Stanford Study: Law Professors Prefer AI Answers Over Peers 75% of the Time

A Stanford Law School study led by Professor Julian Nyarko found that law professors overwhelmingly prefer AI-generated answers to student questions over responses written by fellow instructors. In a blind evaluation of nearly 3,000 anonymized comparisons across 16 U.S. law schools, AI responses won 75% of head-to-head matchups against peer-written answers.
Study Design & Results
The study, titled Law Professors Prefer AI Over Peer Answers, focused on contract law. Participants created 40 representative questions that students might ask after class or during office hours. Professors wrote their own answers, then evaluated responses without knowing whether they came from AI or other professors. The AI systems performed comparably to the best human instructor in the study.
Key findings:
- AI won 75% of head-to-head comparisons against peer answers
- AI responses flagged as pedagogically harmful only 3.5% of the time
- Peer-written answers flagged as harmful 12% of the time
- Evaluations focused on nuanced legal reasoning, not factual recall
Implications for Legal Education
“This study challenges important assumptions about AI’s role in legal education,” Nyarko said. “We focused on law precisely because it requires judgment, nuanced reasoning, and the ability to navigate ambiguity—not just factual recall.”
The research also examined specific AI models including commercial tutoring systems and Google’s NotebookLM, finding varying levels of performance. Even when context limitations affected AI responses, professors still frequently preferred them to human-written alternatives.
Co-author Sarath Sanga from Yale Law School noted: “In most fields where AI gets tested, there’s a right answer. In law, there often isn’t. Two opposing arguments can both be good.”
The study is particularly notable because previous AI evaluations focused on subjects with clear right-or-wrong answers, whereas legal reasoning demands careful analysis of competing arguments and defensible conclusions.
Cautions & Open Questions
Nyarko cautioned against wholesale adoption: “How to implement these tools to most effectively improve student learning is still an open question.” The study evaluated answer quality but noted that implementation challenges such as hallucinations, overreliance, and erosion of critical thinking skills remain.
📖 Read the full source: HN AI Agents
👀 See Also

SWE-rebench Leaderboard Update: February 2026 Results Show Tight Competition
The SWE-rebench leaderboard has been updated with February 2026 results testing 57 fresh GitHub PR tasks. Claude Opus 4.6 leads with 65.3% resolved rate, but the top six models are within 5 percentage points.

Claude-Code v2.1.110 adds TUI mode, push notifications, and multiple fixes
Claude-Code v2.1.110 introduces a new /tui command for flicker-free rendering, push notification capabilities for mobile alerts, and improvements to plugin management and remote control functionality. The release also includes numerous bug fixes for MCP servers, session handling, and UI issues.

Current State of Chinese LLMs: Market Leaders, Open Models, and Business Models
A Reddit analysis details the Chinese LLM landscape, identifying ByteDance's Doubao as the proprietary market leader and DeepSeek as the most innovative, while outlining the business models of major players and 'Six AI Small Tigers' focused on open-weight models.

Claude Code v2.1.132: SIGINT Graceful Shutdown, MCP Fixes, and Terminal Handling Overhaul
Claude Code v2.1.132 fixes graceful shutdown on external SIGINT, adds CLAUDE_CODE_SESSION_ID and CLAUDE_CODE_DISABLE_ALTERNATE_SCREEN env vars, patches MCP memory leaks and tool listing retries, and resolves dozens of terminal edge cases across IDE terminals.