JetBrains on Building a RAG Pipeline for Semantic Code Search: Parsing, Chunking, Vectorization
JetBrains published the first part of a developer diary on building Air Context, their semantic code search platform, a RAG pipeline that feeds LLM agents "precise, citable evidence from real repositories instead of whatever grep happens to surface." Written by Adam Malek and Ashot Kazaryan, Part 1 covers parsing, chunking, and vectorization.
Why semantic search, not grep
The post argues that keyword search and grep fail because they require the agent to know the exact text in advance. The concrete example: "an agent looking for where session tokens get refreshed cannot rely on the code helpfully containing the word 'refresh'." To reason over abstract domains, the agent needs to search by meaning. RAG indexes source code in a way that captures semantics, then lets the agent retrieve relevant pieces on demand via free-text search.
Parsing and chunking
Parsing/chunking is the pre-processing step, and JetBrains calls it "often overlooked." Production repos contain thousands of files spanning hundreds or thousands of lines each, and agents inflate codebases further by being "prolific writers." Files may contain many classes, fields, and methods with varying relatedness.
The wrong approaches, per the post:
- Whole-file embedding — even if you could fit the file into the embedding model, search would return the entire file, defeating the goal of finding a specific function, symbol, or snippet.
- Per-line embedding — individual lines are "semantically insignificant without the surrounding context." A generic function name or comment "does not merit embedding and will produce the wrong retrieval result," overloading the agent with insignificant micro-results.
The goal is properly scoped units: the agent's exploration workflow is mostly concerned with finding a specific function, symbol, or code snippet, so chunk boundaries should align with those units (the "fine AST of parsing and chunking," in the post's phrasing).
What comes next
This is Part 1 of a series. The authors say they'll cover each stage from pre-processing to storage and agent integration, including the "wrong turns" they took. The piece stops mid-sentence before detailing the vectorization stage specifics, so expect a follow-up on how chunks are transformed into an embedding representation.
Who it's for
Developers building retrieval for coding agents on large codebases, or anyone evaluating chunking strategies for code RAG.
📖 Read the full source: HN LLM Tools
👀 See Also

Anthropic Open-Sources Claude for Legal: Plugin Suite for Contract Review, NDA Triage, and More
Anthropic released Claude for Legal, a repo of plugins, agents, and MCP connectors for legal workflows including vendor agreement review, NDA triage, and regulatory monitoring.

Task-observer: A Meta-Skill That Automates Skill Improvement for AI Coding Agents
Task-observer is a meta-skill that self-improves all your AI agent's skills, including itself. It logged 600 skill improvements across 40 skills in 3 months and automates skill creation from work gaps.

Testing MiniMax M2.7 via API on Three Real ML and Coding Workflows
A developer benchmarks MiniMax M2.7 against Claude Opus 4.7 on three real tasks: refactoring a PyTorch project, drafting Obsidian notes, and more. Key findings and setup included.

Claude Code Plugin Yoink Replaces Library Dependencies to Reduce Supply Chain Risk
Yoink is a Claude Code plugin that removes complex dependencies by reimplementing only needed functions, using a three-step workflow with /setup, /curate-tests, and /decompose commands. It currently supports Python with TypeScript and Rust support underway.