AI Subroutines: Deterministic Browser Automation with Zero Token Cost

What AI Subroutines Do
AI Subroutines record browser tasks once and save them as callable tools that replay at zero token cost, zero LLM inference delay, and with 100% determinism. The generated script executes inside the webpage itself, not through a proxy, headless worker, or out-of-process solution.
Key Architectural Decision
The script executes inside the webpage's execution context, which means all authentication, CSRF tokens, TLS sessions, and signed headers get added to requests automatically. No certificate installation, TLS fingerprint modification, or separate auth stack maintenance is required.
Recording Mechanism
During recording, the extension intercepts network requests using two layers:
- MAIN-world fetch/XHR patch installed before any page script runs
- Chrome's webRequest API as a correlated fallback for CORS and service-worker paths
Request bodies including FormData, Blob, and raw bytes are captured, not just JSON.
Network Capture Processing
The system scores and trims approximately 300 requests down to about 5 based on multiple signals:
- First-party vs. third-party origin (+20 / −15)
- Known telemetry hosts (Sentry, Segment, Hotjar, RUM): −80
- Temporal correlation to DOM events (+28 within 800ms, +16 within 2.5s)
- Method and payload shape (mutating POST/PUT/PATCH/DELETE: +35; GET: +5; with request body: +8)
- Response quality (2xx: +12; 4xx+: −25; non-empty body: +4)
- Volatile operation identifiers (−18) for GraphQL queryId, doc_id, operationHash
Volatile GraphQL operation IDs trigger a DOM-only fallback before they break silently on the next run.
Generated Code Structure
The generated code combines network calls with DOM actions (click, type, find) in the same function via an rtrvr.* helper namespace. The top five ranked requests plus DOM interactions get rendered into a 12,000-character context for the generator.
Usage Pattern
Point an AI agent at a spreadsheet of 500 rows, and with just one LLM call, parameters are assigned and 500 Subroutines are kicked off.
Key Use Cases
- Record sending an Instagram DM, then have a reusable routine to send DMs at zero token cost
- Create a routine to get latest products in a site catalog, call it to get thousands of products via direct GraphQL queries
- Set up a routine to file EHR forms based on parameters, with AI inferring parameters from current page context
- Reuse routines daily to sync outbound messages on LinkedIn/Slack/Gmail to a CRM using an MCP server
Why This Matters
The fundamental problem with browser agents for repetitive tasks is that going through the inference loop is unnecessary. Recording once and having the LLM generate a script that leverages all possible interaction methods (direct API calls, DOM interactions, third-party tools/APIs/MCP servers) provides deterministic, cost-effective automation.
📖 Read the full source: HN LLM Tools
👀 See Also

Prism MCP v5.1 adds 10x memory compression and agent learning from corrections
Prism MCP v5.1 introduces 10x memory compression via TurboQuant ported to TypeScript, enabling millions of memories on a laptop without vector databases. The update adds agent learning from user corrections and a visual knowledge graph interface.

Bifrost LLM Gateway: 11 Microsecond Overhead, Single Binary in Go
Bifrost is an open-source LLM proxy written in Go that routes requests to OpenAI, Anthropic, Azure, and Bedrock with 11 microsecond overhead per request, handling 5,000 RPS on a $20/month VPS.

Claude Code HUD: Terminal Dashboard for Monitoring AI Coding Sessions
claude-code-hud is a terminal dashboard that provides real-time monitoring for Claude Code sessions, showing context window usage, API rate limits, and file changes without requiring an IDE. Run it with npx claude-code-hud.

Local Behavioral Monitoring System with MCP Pipeline and Claude Code
A developer built a local behavioral monitoring system called BRAIN that tracks app switches, file operations, and dev sessions, piping data through a custom MCP server to Claude Code. The system runs 100% locally with zero cloud dependency.