Claude-Real-Video: Scene-Aware Frame Extraction + Transcript for Any LLM

Claude-real-video (GitHub) is a Python CLI tool that lets any LLM actually “watch” a video by extracting scene-aware, deduplicated frames and a transcript — all locally, no uploading to external services. Unlike fixed-interval frame sampling (e.g., 1 fps) that over-samples static content and misses fast cuts, this tool uses scene-change detection plus a density floor to capture every meaningful visual change while discarding near-duplicates.
Key Features
- Scene-change detection with configurable sensitivity (
--scene 0.30) - Sliding-window dedup (
--dedup-window 4,--dedup-threshold 8%) — avoids resending the same shot after a cutaway - Density floor (
--fps-floor 1.0) ensures at least one frame every N seconds - Hard cap on total frames:
--max-frames 150 - Whisper transcription with language detection (
--lang autoor specifyen,zh) - Supports URLs (YouTube, Instagram, TikTok, etc.) via yt-dlp and local files
- Outputs a clean folder:
crv-out/frames/*.jpg,crv-out/transcript.txt,crv-out/MANIFEST.txt— drop these into Claude, ChatGPT, or Gemini - Option to generate a visual report of keep/drop decisions (
--report)
Installation
pip install claude-real-video # core (frames + dedup)
pip install "claude-real-video[whisper]" # + audio transcription
System requirement: ffmpeg (install via brew install ffmpeg, sudo apt install ffmpeg, or winget install Gyan.FFmpeg).
Usage Examples
# YouTube/Instagram link
crv "https://www.youtube.com/watch?v=..."
Local file with English transcript
crv lecture.mp4 -o out --lang en
Frames only, no transcript
crv clip.mp4 --no-transcribe
With cookie file for login-gated content
crv "https://..." --cookies cookies.txt
How It Works
The tool uses yt-dlp to fetch URL-based videos (with optional cookies), then ffmpeg extract passes to grab frames at every scene change plus a density floor. A sliding-window near-duplicate detector collapses repeated shots. Audio is transcribed via the Whisper CLI. Everything stays local — no upload to any cloud service.
📖 Read the full source: HN AI Agents
👀 See Also

Building a Sub-500ms Voice Agent: Architecture and Performance Insights
A developer built a voice agent from scratch achieving ~400ms end-to-end latency with full STT → LLM → TTS streaming. Key insights include treating voice as a turn-taking problem, using semantic end-of-turn detection, and colocating all components for minimal latency.

SimplePDF Copilot: Client-Side AI Tool Calling for PDF Form Filling
SimplePDF Copilot uses client-side tool calling to let an LLM fill fields, add fields, delete pages, and more in PDFs — without the PDF leaving the browser.

SOPHIA Meta-Agent for AI Agent Maintenance
SOPHIA is a meta-agent designed as a Chief Learning Officer that observes, diagnoses, researches, and proposes improvements to other AI agents in production ecosystems. The system was designed through 7 iterations using 4 frontier models with human approval required for all deployments.

Claude Code v2.1.76 System Prompt Updates: Security Monitor Refinements and New Hook Event
Claude Code v2.1.76 includes updates to system prompts with 43 new tokens, featuring refinements to the security monitor for autonomous agents and the addition of a PostCompact hook event. Changes include clarified sensitive data detection, expanded code deserialization examples, and improved formatting for irreversible local destruction guidance.