All news
aiproductsaas

claude-real-video: An Open-Source Fix for Vision LLM Limits

08 Aug 2026

The problem: vision LLMs can't watch full videos

Vision-enabled LLMs are increasingly used to analyze video content, but they face a hard constraint: a realistic frame budget of about 150 images per video. That's the working assumption behind a new open-source tool called claude-real-video, a local video processing pipeline built to make the most of that limited budget by intelligently selecting which frames actually matter.

The tool is MIT-licensed and available on GitHub (github.com/HUANGCHIHHUNGLeo/claude-real-video), with an installable MCP server option via pip install 'claude-real-video[mcp]'.

How it works

Rather than feeding every frame to a vision LLM, claude-real-video uses a deduplication process to filter out frames that don't add new information:

  • A 16×16 RGB signature is used by the dedup comparator to summarize each frame.
  • Frames are dropped if fewer than 8% of cells changed between comparisons.
  • A cell counts as "changed" if it moved more than 25/255 in any color channel.
  • A separate local state changes channel operates on a finer 192×192 signature to catch more granular shifts.
  • Frames that pass these checks are resized to 768px on the long side before being sent inline to the vision model.

The goal is to compress a long or complex video down to the ~150 frames that best represent meaningful change, without blowing through cost or context-window constraints.

Testing surfaced an unresolved issue

The author ran the tool on 2,181 real videos, and during that process a user found an issue. The report does not specify what the issue was or whether it has since been resolved — a gap worth watching for anyone considering the tool for production use.

Risks and open questions

Several aspects of the pipeline remain unaddressed publicly:

  • The fixed ~150-image budget could cause important moments in longer or more complex videos to be missed entirely.
  • The fixed dedup thresholds (8% cell change, 25/255 movement) may cause the pipeline to drop frames with subtle but meaningful changes — for example, small but important visual cues that don't trigger the movement threshold.
  • The unresolved issue found during the 2,181-video test run may point to edge cases not yet fully handled.

Additionally, there are no published performance benchmarks, accuracy comparisons, or evaluation results for the dedup or state-change approach, and no details on how the 8% and 25/255 thresholds were chosen. The report also doesn't clarify why 150 images is considered the "realistic" budget — including which model(s) this figure applies to or whether it's driven by cost, context-window size, or both.

Why founders should care

For founders building video-understanding features on top of vision LLMs, this pipeline is likely to be a useful reference point rather than a plug-and-play solution:

  • Teams designing video-LLM products should probably plan around strict frame-selection limits (roughly 150 images) rather than assuming they can feed models unlimited footage.
  • The dedup and resizing techniques may offer a low-cost, practical pattern worth adapting — using signature comparisons and movement thresholds to cut frame counts before sending data to a model.
  • The issue that surfaced during large-scale testing suggests that similar video pipelines likely require extensive real-world testing before founders rely on them in production, since edge cases may not be fully worked out even after testing on thousands of videos.
  • Because the tool is MIT-licensed and available via pip as an MCP server, it could serve as a drop-in component for startups prototyping video-analysis agents — potentially saving early engineering time, though founders should independently validate accuracy and edge-case handling before shipping.

Given the missing benchmarks and unresolved test issue, teams evaluating claude-real-video should treat it as an early-stage building block — promising for prototyping, but likely needing further validation before it underpins a production video-analysis feature.

Sources