Video-Use

An agent-driven video editing skill from the team behind browser-use. The model never watches the video: it reads it through a word-level timestamped transcript with speaker identification and audio event detection, which lets it cut on a word boundary rather than on an approximate second. Filler word removal, automatic colour grading, audio fades, burned-in subtitles and overlay animation generation through HyperFrames, Remotion, Manim or PIL. It installs as a Claude Code skill by cloning then symlinking, and relies on ffmpeg, plus yt-dlp when the source is an online address.

Strengths

Limitations

Best for

View on Coeurdar