Measuring the Sloppiness of Code: Verbosity and Erosion Metrics
LLMs have gotten very good at generating code that passes tests. But correct code can still be sloppy: full of duplicated lines, unnecessary abstractions, and poor design choices. As the author notes, "just because the code is formally correct doesn't mean that it is not introducing unnecessary abstractions, creating duplicates, or just making bad decisions overall." The result is an explosion in lines of code (LOC) that's hard for humans to keep up with, and—counter to some claims—agents can't manage the slop either.
Why LLM-as-judge doesn't work
The article dismisses the common industry approach of using AI to judge code quality. Asking a model to rate code 1–10 is "basically equivalent to a random number generator." Pairwise comparisons (A vs. B) are unstable: simply renaming the solutions can flip the model's preference. Rubrics and LLM-written tests help but are "still a far shot from actually getting rid of the slop."
The simplest metric: change in LOC
Surprisingly effective: just tracking the change in the number of lines of code. The author notes the irony: "if we started optimizing for it, it would cease to be a meaningful measure."
Verbosity
Measures duplicated and unnecessarily verbose lines. It is the fraction of lines flagged by AST-Grep or marked as clones, divided by total LOC:
Verbosity = |AST-Grep flagged lines ∪ clone lines| / LOC
Erosion
Measures how much of a codebase's mass is concentrated in a few large, complex functions. First define the mass of a function:
mass(f) = CC(f) * sqrt(SLOC(f))
Where CC(f) is cyclomatic complexity and SLOC(f) is source lines of code. Then erosion is the fraction of total mass held by functions with cyclomatic complexity greater than 10:
Erosion = ∑_{f: CC(f) > 10} mass(f) / ∑_f mass(f)These two measures—introduced by the paper SlopCodeBench—separated legacy codebases from LLM-generated slop quite well in the author's tests.
Takeaways
- LLM judges are unreliable for code quality; preference flips on renaming.
- Change in LOC is a surprisingly strong sloppiness signal, but breaks if optimized directly.
- Verbosity combines AST-Grep flags and clone detection over LOC.
- Erosion captures complexity mass in functions with CC > 10.
For teams shipping LLM-generated code at scale, these metrics offer a quantitative alternative to vibes-based evaluation.
📖 Read the full source: HN LLM Tools
👀 See Also

Qwen3.5-122B-A10B-MINT-MLX runs smoothly on M5 Pro with 64GB RAM
A user reports successful local deployment of the Qwen3.5-122B-A10B-MINT-MLX model on an M5 Pro with 64GB RAM, achieving 39.58 tokens/sec generation speed with specific VRAM allocation commands.

Research on AI Agent Consistency: Key Findings and Practical Takeaways
A study of 3,000 experiments across Claude, GPT-4o, and Llama reveals that consistent agents achieve 80–92% accuracy while inconsistent ones drop to 25–60%, with 69% of divergence occurring at the first tool call.

Research on Professional Social Networks for AI Agents
Analysis of intent, behavior, and platform trends for professional AI agent social networks, focusing on Moltbook, Agent.ai, and Clawsphere, with examination of Meta's acquisition impact.

Bird Skill Repository Removed — Backup Your X/Twitter Access Now
The popular bird skill by @steipete has been removed from GitHub. Users should backup their installations immediately.