Video-talkcraft Turns Claude Code Into a Frame-Perfect Video Director

A new open-source agent skill hands Claude Code and Codex a full motion-design studio for voiceover explainer videos with word-level sync and Remotion rendering.

·
·
Video-talkcraft Turns Claude Code Into a Frame-Perfect Video DirectorPRO
  • New video-talkcraft agent skill turns Claude Code/Codex into an explainer-video motion studio.
  • Ships 78 motion recipe cards with self-contained Remotion tsx source and runnable HTML previews.
  • Word-level voiceover alignment hits 20-40ms median deviation on 110s mixed-language audio.
  • Seven-layer camera system makes static frames structurally impossible, with automatic still-frame detection.
  • Triple QA loop: still detection, per-cue SFX energy check, and motion-anchor frame review.
  • Free for personal, educational, and research use under PolyForm Noncommercial 1.0.0.

Video-talkcraft is an open-source agent skill for Claude Code and Codex that produces voiceover-driven explainer videos through Remotion. Hand it a script and a finished audio file, and it aligns word-level timestamps, writes a shot-by-shot storyboard, and renders the final piece with kinetic captions, camera moves, transitions, and sound effects locked to the narration.

Sync, not slideshow

Most AI video tools stitch stock footage together or animate a slide deck. Video-talkcraft targets something narrower: narrated explainers where every motion beat lands on a specific syllable. The agent reads a methodology document, selects from 78 motion recipe cards, writes Remotion code, and runs a triple-QA loop before delivering the output.

What ships in the repo

  • 78 motion recipe cards. Each card includes intent, parameters, known pitfalls, self-contained Remotion TSX source, and a runnable HTML preview. A gallery page shows all 78 previews at once.
  • Word-level voiceover alignment. A CPU script aligns narration to audio using FireRedASR2-CTC int8 by default, with faster-whisper as a fallback that skips manual downloads. On a 110-second Chinese-English mixed voiceover, character-level deviation clocked in at a median of 20–40ms, worst case 200ms, with zero false positives in QA.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads