xAI Ships Grok Imagine Video 1.5 With Native Audio and 2x Faster Generation
xAI's Grok Imagine Video 1.5 exits preview with native audio, faster generation, and the top spot on the image-to-video leaderboard
- GA release: Grok Imagine Video 1.5 is now generally available via the xAI API as
grok-imagine-video-1.5and on grok.com/imagine. - Leaderboard #1: Scored 1473 Elo on the Image-to-Video Arena, a +52 point jump over v1.0, beating Seedance 2.0 and Google Veo.
- Native audio in one pass: Sound effects, ambience, lip-synced dialogue, and background music are generated alongside the video with no separate audio step.
- Speed boost: The Fast variant produces 6-second 720p clips in ~25 seconds, nearly double the speed of the previous model.
- Pricing: API billed per second of output -- $0.08/sec at 480p and $0.14/sec at 720p (a 15-second 720p clip costs ~$2.10).
- Known limits: Quality degrades after 2-3 chained clip extensions; max resolution is 720p with 1080p on the roadmap.
xAI just pushed Grok Imagine Video 1.5 out of preview and into general availability, and the timing is deliberate. The announcement dropped alongside a cinematic trailer for a film called Odyssey, produced entirely with the model by creator David Thompson (@heavypulp). The message is clear: this is not a research demo. It is a production tool.
From still image to cinematic clip
The model's core job is straightforward: give it a starting frame and a prompt describing the motion, and it animates the scene, including camera moves, atmosphere, and physics, while staying faithful to your source image. Grok Imagine Video 1.5 is not trying to invent the entire scene from scratch. It starts from your image, then animates it. That distinction matters enormously for production workflows where visual identity is already locked in.
The model supports three generation modes: text-to-video from a prompt alone, image-to-video that animates a still input, and reference-to-video that grounds the output in up to seven reference images for consistent characters, styles, or settings. It produces short videos between 1 and 15 seconds at 24 fps, available at 480p or 720p across seven aspect ratios.
The headline upgrades in 1.5
Compared to the previous model, 1.5 improves across every dimension that matters for real creative work: sound effects, ambience, and dialogue are generated in the same pass and land on the action; speech is clearer and better synced; and movement holds together over the length of a clip with fewer warps and more believable weight and momentum.
- Native audio in one pass: A single generation can include background music, sound effects, and lip-synced dialogue, so you do not need a separate audio pass after the clip is rendered.
- Speed: Grok Imagine Video 1.5 Fast almost doubles generation speed, producing 6-second, 720p videos in about 25 seconds, down from 40+ seconds in the previous model.
- Motion coherence: The model is built on Aurora, xAI's autoregressive architecture, which minimizes character warping and maintains visual consistency across camera changes and scene transitions without requiring manual intervention.
- Clip extension: The native Extend from Frame feature lets you add 6 to 10 seconds per extension, building longer sequences from your initial clip without re-generating from scratch.
- Parallel generation: You can kick off multiple agents in parallel on your projects, so instead of waiting for one generation to finish before starting the next, you can run several prompts at once.
The architecture underneath
Aurora adopts a structure called an "autoregressive mixture-of-experts,