Claude-Real-Video: Scene-Aware Frame Extraction + Transcript for Any LLM

✍️ OpenClawRadar📅 Published: July 3, 2026🔗 Source
Claude-Real-Video: Scene-Aware Frame Extraction + Transcript for Any LLM
Ad

Claude-real-video (GitHub) is a Python CLI tool that lets any LLM actually “watch” a video by extracting scene-aware, deduplicated frames and a transcript — all locally, no uploading to external services. Unlike fixed-interval frame sampling (e.g., 1 fps) that over-samples static content and misses fast cuts, this tool uses scene-change detection plus a density floor to capture every meaningful visual change while discarding near-duplicates.

Key Features

  • Scene-change detection with configurable sensitivity (--scene 0.30)
  • Sliding-window dedup (--dedup-window 4, --dedup-threshold 8%) — avoids resending the same shot after a cutaway
  • Density floor (--fps-floor 1.0) ensures at least one frame every N seconds
  • Hard cap on total frames: --max-frames 150
  • Whisper transcription with language detection (--lang auto or specify en, zh)
  • Supports URLs (YouTube, Instagram, TikTok, etc.) via yt-dlp and local files
  • Outputs a clean folder: crv-out/frames/*.jpg, crv-out/transcript.txt, crv-out/MANIFEST.txt — drop these into Claude, ChatGPT, or Gemini
  • Option to generate a visual report of keep/drop decisions (--report)
Ad

Installation

pip install claude-real-video           # core (frames + dedup)
pip install "claude-real-video[whisper]"  # + audio transcription

System requirement: ffmpeg (install via brew install ffmpeg, sudo apt install ffmpeg, or winget install Gyan.FFmpeg).

Usage Examples

# YouTube/Instagram link
crv "https://www.youtube.com/watch?v=..."

Local file with English transcript

crv lecture.mp4 -o out --lang en

Frames only, no transcript

crv clip.mp4 --no-transcribe

With cookie file for login-gated content

crv "https://..." --cookies cookies.txt

How It Works

The tool uses yt-dlp to fetch URL-based videos (with optional cookies), then ffmpeg extract passes to grab frames at every scene change plus a density floor. A sliding-window near-duplicate detector collapses repeated shots. Audio is transcribed via the Whisper CLI. Everything stays local — no upload to any cloud service.

📖 Read the full source: HN AI Agents

Ad

👀 See Also

Building a Sub-500ms Voice Agent: Architecture and Performance Insights
Tools

Building a Sub-500ms Voice Agent: Architecture and Performance Insights

A developer built a voice agent from scratch achieving ~400ms end-to-end latency with full STT → LLM → TTS streaming. Key insights include treating voice as a turn-taking problem, using semantic end-of-turn detection, and colocating all components for minimal latency.

OpenClawRadar
SimplePDF Copilot: Client-Side AI Tool Calling for PDF Form Filling
Tools

SimplePDF Copilot: Client-Side AI Tool Calling for PDF Form Filling

SimplePDF Copilot uses client-side tool calling to let an LLM fill fields, add fields, delete pages, and more in PDFs — without the PDF leaving the browser.

OpenClawRadar
SOPHIA Meta-Agent for AI Agent Maintenance
Tools

SOPHIA Meta-Agent for AI Agent Maintenance

SOPHIA is a meta-agent designed as a Chief Learning Officer that observes, diagnoses, researches, and proposes improvements to other AI agents in production ecosystems. The system was designed through 7 iterations using 4 frontier models with human approval required for all deployments.

OpenClawRadar
Claude Code v2.1.76 System Prompt Updates: Security Monitor Refinements and New Hook Event
Tools

Claude Code v2.1.76 System Prompt Updates: Security Monitor Refinements and New Hook Event

Claude Code v2.1.76 includes updates to system prompts with 43 new tokens, featuring refinements to the security monitor for autonomous agents and the addition of a PostCompact hook event. Changes include clarified sensitive data detection, expanded code deserialization examples, and improved formatting for irreversible local destruction guidance.

OpenClawRadar