Video-Use
An agent-driven video editing skill from the team behind browser-use. The model never watches the video: it reads it through a word-level timestamped transcript with speaker identification and audio event detection, which lets it cut on a word boundary rather than on an approximate second. Filler word removal, automatic colour grading, audio fades, burned-in subtitles and overlay animation generation through HyperFrames, Remotion, Manim or PIL. It installs as a Claude Code skill by cloning then symlinking, and relies on ffmpeg, plus yt-dlp when the source is an online address.
Strengths
- Cuts on a word boundary thanks to a timestamped transcript, not on an approximate second
- The edit is described in plain language, with no video editor interface to learn
- Footage stays on your machine, only the audio leaves for transcription
Limitations
- An ElevenLabs key is mandatory and billed per source: the repository is open, the usage is not
- Install by clone and symlink, with ffmpeg to set up yourself: nothing zero-install about it
- The agent decides cuts from text, so it misses what is not spoken: a look, a gesture, a shot that lingers
Best for
- Producing short formats in series that all follow the same editing template
- Cleaning up a spoken take without opening an editor: filler words, gaps, subtitles
- Rescuing an over-long take by cutting from what is said in it, when eyeballing the timeline would take the evening