Infracost cuts Claude token usage 79% by redesigning CLI for AI agents

✍️ OpenClawRadar📅 Published: May 19, 2026🔗 Source
Infracost cuts Claude token usage 79% by redesigning CLI for AI agents
Ad

Infracost, a CLI tool that estimates cloud infrastructure costs from Terraform, CloudFormation, and CDK, has redesigned its output for AI coding agents like Claude Code and Cursor. The result: up to 79% fewer output tokens and 67% lower API costs vs a bare-Claude baseline. The redesign revolves around two techniques: predicate pushdown into the CLI and a token-efficient output format.

Benchmark details

  • 16 questions over a 3-project Terraform fixture with 1,171 resources
  • Model: Claude Opus, 5 repeats per question
  • Baseline: bare Claude with Bash and Read tools, no skill loaded
  • Compared against Infracost skill with --llm output flag

Key results

MetricBare ClaudeWith Infracost skill (--llm)Change
Correct answers5 / 11 (45%)11 / 11 (100%)+6
Total cost (USD)$16.41$9.63-41%
Output tokens207,01781,697-61%
Wall time50 min50 mintied

One example: the question "count distinct resources failing the tagging policy, deduplicated across projects" cost $3.51 with bare Claude and hit the 25-turn cap, returning no answer. With the redesigned CLI, the same question cost $0.25 and returned the correct answer.

Ad

Technical approach

  • Predicate pushdown: Instead of having the agent pipe JSON through jq or write Python parsers, the CLI accepts filtering flags (e.g., --tag-policy), offloading computation to the tool itself. This reduces the number of turns and token consumption.
  • Token-efficient output format: The --llm flag returns a compact, agent-friendly format rather than verbose human-readable tables or full JSON. This alone accounts for a significant share of the reduction.

Benchmark harness gotchas

Infracost open-sourced their harness setup to help others avoid pitfalls:

  • Sandbox HOME for baseline runs to avoid accidental skill loading
  • Set TMPDIR to a project-local directory to circumvent macOS ACL issues
  • Prepend the test binary to PATH rather than relying on system install
  • Use 5+ repeats per cell due to 20-30% token variance
  • Re-run cells that hit turn caps (--rerun-failed) and re-score if the verifier changes (--rescore)

If you maintain a CLI that AI agents call as a subprocess, the same two moves — predicate pushdown and a dedicated agent output format — likely apply. The redesign also improved the human-facing CLI, though the article focuses on the agent path.

📖 Read the full source: HN AI Agents

Ad

👀 See Also

LM Studio plugins add web image analysis for vision-capable LLMs
Tools

LM Studio plugins add web image analysis for vision-capable LLMs

A developer created plugins for LM Studio that enable vision-capable LLMs to fetch and analyze images from the web, with automatic image processing and tool chaining. The plugins work with models like Qwen 3.5 9b/27b and include updated Duck-Duck-Go and Visit Website functionality.

OpenClawRadar
🦀
Tools

Researcher Builds Veracity-Checking Skill for Claude Code, Finds Hallucinations in Own Documentation

A researcher built a Claude Code skill called /veracity-tweaked-555 that decomposes documents into atomic claims and verifies each via web search using 16 parallel agents across 4 waves. When self-audited, the skill scored 62/100 due to fabricated statistics and inflated claims in its own documentation.

OpenClawRadar
Claude Code CLI Toolkit: Four Tools for Code Review, Project Briefs, Auto-Journaling Git Hooks
Tools

Claude Code CLI Toolkit: Four Tools for Code Review, Project Briefs, Auto-Journaling Git Hooks

A developer has released four CLI tools built around Claude Code's print mode that handle code reviews, project brief generation, auto-journaling git hooks, and Claude session status. The tools use existing Claude Code authentication and are available as open source.

OpenClawRadar
Local semantic search for AI conversations with fastembed and LanceDB
Tools

Local semantic search for AI conversations with fastembed and LanceDB

A developer indexed 368K AI conversation messages locally using fastembed for CPU-based embeddings and LanceDB as a serverless vector store, achieving 12ms p50 search latency without API keys.

OpenClawRadar