TREX: Greptile's AI Code Reviewer That Runs Your Code

Greptile released TREX (Test, Run, Execute), an execution layer that runs your code during AI-powered code review. Instead of just reading diffs, TREX actually executes the changed code and surfaces runtime bugs — UI regressions, state-dependent logic errors, race conditions — that static analysis can't catch.
Architecture: Orchestrator + Per-Issue Subagents
Early versions tried separate agents or a single combined agent. Both failed: separate agents duplicated work with no shared context; a single agent got overloaded managing setup, screenshots, and tests. The solution was an orchestrator agent (the main Greptile reviewer) that reads the diff, identifies suspicious issues, and spins up a dedicated TREX subagent per issue, all running in parallel. Each subagent inherits the orchestrator's context and has its own context window scoped to its specific investigation.
Example: a UI feature behind an auth gate. A subagent autonomously sets up the environment, handles authentication, toggles feature flags, and returns a screenshot of the rendered feature.
Multi-Modal Artifacts vs. Bullet Points
Initial TREX output was bullet-point summaries — but bullet points allowed hallucinations (e.g., claiming a test passed when it hadn't) and gave no way to verify. The fix: each TREX finding is backed by a set of multi-modal artifacts: screenshots, execution logs, API traces, and execution scripts. Every modality tells part of the story, making it possible to trace exactly what happened. The first artifact that impressed the team was a video capture of an animation change — showing the actual runtime effect.
What It Catches
TREX targets bugs that don't appear in code diffs: logic errors requiring specific state sequences, UI regressions after page load, and race conditions that need real requests. It generates and runs tests, but the focus is on finding bugs, not just writing tests. The subagent figures out setup on its own.
As Shlok Mehrotra, the engineer behind TREX, puts it: "You can read the diff perfectly and still miss these types of bugs completely."
📖 Read the full source: HN AI Agents
👀 See Also

Codebook Lossless LLM Compression: 10-25% RAM Reduction with Bitwise Packing
A developer's proof-of-concept code demonstrates lossless LLM compression by packing fp16 weights into blocks, achieving 10-25% RAM reduction with a trade-off of approximately halved inference speed. The approach identifies that most models only use 12-13 bits of unique values despite fp16's 16-bit representation.

Developer builds local AI research agent that creates podcasts from topics or YouTube links
A developer built a fully local AI agent that takes topics or YouTube links and generates deep-dive reports, conversational podcast scripts, and audio. The system dynamically researches, extracts insights, refines summaries, and creates natural back-and-forth conversations.

Toothcomb: Open-Source Real-Time Speech Fact-Checker Built with Claude Opus and Sonnet APIs
Toothcomb is an open-source tool that takes a speech transcript, fact-checks claims, detects logical fallacies and manipulative language using Claude Opus API, and supports real-time microphone streaming.

Running Multiple Claude Code Sessions in Parallel with Git Worktrees
A developer shares how they use git worktrees to run multiple Claude Code sessions on separate branches without stashing or context switching. Review diffs, merge, and move on.