Persistent Indexes Over Extraction: Architecture for a YouTube MCP Server

✍️ OpenClawRadar📅 Published: April 15, 2026🔗 Source
Persistent Indexes Over Extraction: Architecture for a YouTube MCP Server
Ad

A developer has shared detailed architecture notes from building a YouTube MCP server that implements persistent local indexes, contrasting with the common "extract-and-forget" pattern observed in over 40 existing servers.

Architecture Decisions

  • Three-tier fallback on every tool: Uses YouTube Data API → yt-dlp → page extraction. Every response includes a provenance field ({sourceTier, fallbackDepth, partial, fetchedAt, sourceNotes}) to prevent silent degradation. Quota exhaustion on tier 1 results in a degraded response with clear provenance instead of a failure.
  • Persistence model: SQLite + sqlite-vec for local vector storage in a single file, with no Docker or external database. Embeddings persist across sessions, allowing knowledge to accumulate—the tenth query on an indexed playlist is richer and faster than the first.
  • Embedding provider abstraction: Uses Gemini text-embedding-004 (768d) when a Gemini key is present, falling back to all-MiniLM-L6-v2 (384d) fully offline via local inference. Both are handled by the same abstraction, enabling semantic search with zero API keys at reduced quality or transparent upgrades when a key is added.
  • Visual search as a separate index: Three independent layers: Apple Vision VNGenerateImageFeatureVectorRequest for per-frame feature prints for image-to-image similarity, Gemini Vision for natural language scene descriptions per keyframe, and Gemini text-embedding-004 for 768d embeddings over OCR text + descriptions for text→visual search. Returns actual frame paths on disk + timestamps + match reasoning, genuinely separate from the transcript pipeline.
  • Token efficiency via strict output schemas: Achieves 75–87% smaller responses than raw YouTube API output by removing thumbnails, eTags, and localization bloat, and using normalized engagement ratios instead of raw counts.
Ad

Tradeoffs Encountered

  • Disk usage grows with persistence: Solved with TTL caches per tool category, a mediaStoreHealth diagnostic, and per-collection cleanup tools.
  • Visual indexing is expensive: Due to keyframe extraction + vision + OCR + embeddings. Made opt-in per video rather than automatic during import.
  • Three-tier fallback adds latency when earlier tiers fail: Considered worth it for reliability, as API quota exhaustion is a real problem in production, and yt-dlp/page extraction keep things working.
  • mcpName vs npm name collision risk: MCP registry uses io.github.<user>/<name> while npm is flat. Solved by making them explicit and different.
  • Apple Vision locks the image-to-image similarity layer to macOS: Accepted tradeoff, as the Gemini-based layers work cross-platform.

The code is open source, and the developer is open to discussing design decisions further, particularly on the persistence vs extraction tradeoff or the visual pipeline.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

PocketTeam: A Claude Code Pipeline with Hook-Based Safety and Learning Agents
Tools

PocketTeam: A Claude Code Pipeline with Hook-Based Safety and Learning Agents

PocketTeam is a Claude Code pipeline that implements 9 safety layers at the tool-call level to block dangerous operations like writes to .env or rm -rf commands. The system includes an Observer agent that analyzes completed tasks and writes structured learnings to improve future agent performance.

OpenClawRadar
claude-real-video: Free Tool to Make Claude Watch Videos with Perception Layer
Tools

claude-real-video: Free Tool to Make Claude Watch Videos with Perception Layer

claude-real-video adds a local perception layer so Claude sees camera moves, pacing, gestures, and voice emotion. Demo watches NVIDIA GTC 2026 keynote and narrates screen action. Free version available.

OpenClawRadar
FixAI Dev: A Consumer Rights Game Using Claude Haiku with Strict JSON Contracts
Tools

FixAI Dev: A Consumer Rights Game Using Claude Haiku with Strict JSON Contracts

A developer built a browser game where Claude Haiku acts as a corporate AI denying consumer requests; players argue using real consumer protection laws across 37 cases in EU, US, UK, and Australia. The architecture uses Haiku for language only, with server-side game logic and strict JSON contracts between components.

OpenClawRadar
Google Research introduces TurboQuant for AI model compression
Tools

Google Research introduces TurboQuant for AI model compression

Google Research has introduced TurboQuant, a compression algorithm that reduces AI model size with zero accuracy loss. It addresses memory overhead in vector quantization and improves key-value cache performance.

OpenClawRadar