Claude Code at Scale: How Agentic Search Avoids RAG Failure Modes in Large Codebases

Claude Code is running in production across multi-million-line monorepos, decades-old legacy systems (C, C++, C#, Java, PHP), and distributed architectures with thousands of developers. Rather than relying on RAG-based retrieval — which fails because embedding pipelines can't keep up with active teams, returning functions renamed two weeks ago or deleted modules — Claude Code navigates codebases like a software engineer: it traverses the file system, reads files, uses grep, and follows references locally without requiring a centralized index to be built, maintained, or uploaded to a server.
The harness matters more than the model
Claude Code's performance is determined less by model benchmarks and more by the harness — five extension points that build on each other:
- CLAUDE.md files — context files loaded automatically at every session start: a root file for the big picture, subdirectory files for local conventions. Keeping them focused on broadly applicable information prevents context-window waste.
- Hooks — not detailed beyond being listed as an extension point.
- Skills — not detailed beyond being listed as an extension point.
- Plugins — not detailed beyond being listed as an extension point.
- MCP servers — not detailed beyond being listed as an extension point.
Two additional capabilities — LSP integrations and subagents — round out the setup. The article advises building these layers in the order listed, as each layer builds on what came before.
Tradeoff: starting context quality
Agentic search works best when Claude has enough starting context to know where to look. Asking it to find all instances of a vague pattern across a billion-line codebase will hit context-window limits before work begins. Teams that invest in codebase setup through CLAUDE.md files see better results.
📖 Read the full source: HN AI Agents
👀 See Also

OpenPlawd: OpenClaw Skill for Automated Plaud Meeting Notes
OpenPlawd is an OpenClaw skill that automatically processes Plaud recordings into structured HTML meeting notes. It polls Plaud accounts hourly, transcribes with Whisper or OpenAI, chunks large files, and generates notes with action items via an OpenClaw agent.

External Reranker Plugin for OpenClaw Memory-Core: Repurpose Old GPUs
A developer shares a plugin for OpenClaw that allows memory-core to use an external reranker instead of the built-in QMD, improving performance on CPU-bound setups.

Optio: Orchestrating AI Coding Agents in Kubernetes from Ticket to PR
Optio is an open-source orchestration system that turns tickets into merged pull requests using AI coding agents like Claude Code or Codex. It handles the full lifecycle in isolated Kubernetes pods with a feedback loop that auto-resumes agents on CI failures or review feedback.

SkyClaw: Rust-Based Autonomous AI Agent Runtime
SkyClaw is an autonomous AI agent runtime built in Rust with a 7.1 MB binary that idles at 14 MB RAM and starts in under one second. It operates on five engineering principles including autonomy, robustness, and brutal efficiency.