Measuring the Sloppiness of Code: Verbosity and Erosion Metrics

✍️ OpenClawRadar📅 Published: September 12, 2026🔗 Source
Ad

LLMs have gotten very good at generating code that passes tests. But correct code can still be sloppy: full of duplicated lines, unnecessary abstractions, and poor design choices. As the author notes, "just because the code is formally correct doesn't mean that it is not introducing unnecessary abstractions, creating duplicates, or just making bad decisions overall." The result is an explosion in lines of code (LOC) that's hard for humans to keep up with, and—counter to some claims—agents can't manage the slop either.

Why LLM-as-judge doesn't work

The article dismisses the common industry approach of using AI to judge code quality. Asking a model to rate code 1–10 is "basically equivalent to a random number generator." Pairwise comparisons (A vs. B) are unstable: simply renaming the solutions can flip the model's preference. Rubrics and LLM-written tests help but are "still a far shot from actually getting rid of the slop."

The simplest metric: change in LOC

Surprisingly effective: just tracking the change in the number of lines of code. The author notes the irony: "if we started optimizing for it, it would cease to be a meaningful measure."

Verbosity

Measures duplicated and unnecessarily verbose lines. It is the fraction of lines flagged by AST-Grep or marked as clones, divided by total LOC:

Verbosity = |AST-Grep flagged lines ∪ clone lines| / LOC
Ad

Erosion

Measures how much of a codebase's mass is concentrated in a few large, complex functions. First define the mass of a function:

mass(f) = CC(f) * sqrt(SLOC(f))

Where CC(f) is cyclomatic complexity and SLOC(f) is source lines of code. Then erosion is the fraction of total mass held by functions with cyclomatic complexity greater than 10:

Erosion = ∑_{f: CC(f) > 10} mass(f) / ∑_f mass(f)

These two measures—introduced by the paper SlopCodeBench—separated legacy codebases from LLM-generated slop quite well in the author's tests.

Takeaways

  • LLM judges are unreliable for code quality; preference flips on renaming.
  • Change in LOC is a surprisingly strong sloppiness signal, but breaks if optimized directly.
  • Verbosity combines AST-Grep flags and clone detection over LOC.
  • Erosion captures complexity mass in functions with CC > 10.

For teams shipping LLM-generated code at scale, these metrics offer a quantitative alternative to vibes-based evaluation.

📖 Read the full source: HN LLM Tools

Ad

👀 See Also