ETH Zurich Study Questions Value of AGENTS.md Files for AI Coding Agents

✍️ OpenClawRadar📅 Published: March 8, 2026🔗 Source
ETH Zurich Study Questions Value of AGENTS.md Files for AI Coding Agents
Ad

Research Findings on AGENTS.md Files

A new paper from ETH Zurich researchers challenges the widespread industry practice of using AGENTS.md files with AI coding agents. The study, conducted by Thibaud Gloaguen, Niels Mündler, Mark Müller, Veselin Raychev, and Martin Vechev, provides empirical evidence that these context files often hinder rather than help AI agents.

Methodology and Testing

The team built AGENTbench, a novel dataset of 138 real-world Python tasks sourced from niche repositories to avoid bias from popular benchmarks like SWE-bench that AI models may have memorized. They tested four agents: Claude 3.5 Sonnet, Codex GPT-5.2, GPT-5.1 mini, and Qwen Code across three scenarios:

  • No context file
  • LLM-generated AGENTS.md file
  • Human-written AGENTS.md file

Performance was measured using three proxy indicators: task success rates (determined by repository unit tests), number of agent steps, and overall inference costs.

Key Results

LLM-generated context files degraded performance, reducing task success rates by an average of 3% compared to providing no context file. These files consistently increased the number of steps agents took, driving up inference costs by over 20%.

Human-written files showed marginal gains with a 4% average increase in task success rate on AGENTbench, but this came with a parallel increase in steps, raising costs by up to 19%.

Including architectural overviews or repository structure explanations in AGENTS.md files did not reduce the time models spent locating relevant files for tasks.

Ad

Behavior Analysis

Trace analysis revealed that agents generally followed instructions in AGENTS.md files, leading them to run more tests, read more files, execute more grep searches, and perform more code-quality checks. While thorough, this behavior was often unnecessary for resolving specific tasks, forcing reasoning models to "think" harder without yielding better final patches.

Practical Recommendations

The researchers recommend omitting LLM-generated context files entirely and limiting human-written instructions to non-inferable details, such as highly specific tooling or custom build commands. They note that while 60,000 open-source repositories currently contain context files like AGENTS.md, and many agent frameworks feature built-in commands to auto-generate them, these files have only marginal effects on agent behavior.

📖 Read the full source: HN AI Agents

Ad

👀 See Also

Pentagon Sends Anthropic Final Offer for Military AI Use Amid Dispute
News

Pentagon Sends Anthropic Final Offer for Military AI Use Amid Dispute

The Pentagon sent Anthropic a best and final offer for unrestricted military use of its Claude AI model, with a Friday deadline to grant full access or face losing military business and being labeled a supply chain risk.

OpenClawRadar
Claude-Code v2.1.31 Release: Key Updates and Bug Fixes
News

Claude-Code v2.1.31 Release: Key Updates and Bug Fixes

Claude-Code v2.1.31 has been released with important enhancements including session resume hints, Japanese IME support, and bug fixes for PDF handling and API requests.

OpenClawRadar
Research shows AI users often accept LLM answers without verification
News

Research shows AI users often accept LLM answers without verification

University of Pennsylvania research found AI users engage in 'cognitive surrender,' accepting LLM answers with minimal scrutiny. In experiments, users accepted correct AI answers 93% of the time and incorrect answers 80% of the time, even when AI was wrong half the time.

OpenClawRadar
Talkie: A 13B LLM Trained Exclusively on Pre-1931 Text, Using Claude as a Judge in RL Training
News

Talkie: A 13B LLM Trained Exclusively on Pre-1931 Text, Using Claude as a Judge in RL Training

Researchers released Talkie, a 13B LLM trained only on text published before 1931 (no internet, no WWII data). Claude Sonnet 4.6 was used as the judge in its online DPO reinforcement learning pipeline, and Claude Opus 4.4 generated synthetic multi-turn conversations for fine-tuning. The model can write Python code from a few in-context examples despite zero modern code in training.

OpenClawRadar