xAI's Grok Imagine Video 1.5 Lite Beats Veo 3.1 at 65% Lower Cost

xAI's distilled video model lands at #17 on the Artificial Analysis leaderboard, beating Veo 3.1 while costing roughly a third as much.

·
·
·
Read6 min
TypeNews
  • xAI released Grok Imagine Video 1.5 Lite, a distilled, cheaper sibling of Grok Imagine Video 1.5.
  • Ranks #17 on AA-Video-T2V v2.0, two places above Google's Veo 3.1.
  • Priced at $0.14/sec at 1080p, roughly a third of Veo 3.1's $0.40/sec.
  • Fastest model at its quality tier: 60.5 seconds median for a 10-second 1080p clip.
  • Supports 1 to 15 second clips, text-to-video and image-to-video, native audio, seven aspect ratios.
  • Available via xAI API, fal, and Vercel AI Gateway.

Grok Imagine Video 1.5 Lite cuts generation costs and latency

xAI has released Grok Imagine Video 1.5 Lite, a distilled version of its flagship video generator. At publication, Artificial Analysis ranks Lite No. 17 on both the AA-Video-T2V v2.0 and AA-Video-T2V-Silent v2.0 leaderboards. It sits two places above Google’s Veo 3.1 while costing about 35% as much at 1080p.

Lite also lands on the benchmark’s Pareto frontiers for quality versus speed and quality versus price. A Pareto-frontier model offers a combination that no measured alternative improves across both dimensions. In this case, the benchmark found no higher-quality option with lower latency or a lower price among the tested configurations.

Distillation drives the trade-off

The OpenRouter model card describes Lite as a faster, lower-cost model distilled from Grok Imagine Video 1.5. Distillation trains a student model to reproduce a larger model’s behavior while reducing inference requirements. Artificial Analysis finds the expected result: lower latency and cost, accompanied by lower scores than the full model in nine of ten capability categories.

Lite supports text-to-video and image-to-video generation, native audio, seven aspect ratios, and clips from 1 to 15 seconds. Available output settings include 480p, 720p, and 1080p, with common formats such as widescreen 16:9 and vertical 9:16.

OpenRouter says the 1080p option renders at 720p and then upscales the result. The output has 1080p frame dimensions, while its native image detail originates from a 720p render. That distinction matters for fine textures, small text, cropping, and compositing.

Pricing changes the iteration budget

xAI prices Lite at $0.14 per second for 1080p output, $0.03 for 720p, and $0.02 for 480p. At 1080p, Lite costs 44% less than the full Grok Imagine Video 1.5 and 65% less than Veo 3.1 with audio.

Published prices in US dollars. The 60-second figures represent aggregate generated footage because Lite limits each clip to 15 seconds.
Model and setting Per second 10 seconds 60 seconds
Grok Imagine Lite, 1080p $0.14 $1.40 $8.40
Grok Imagine Lite, 720p $0.03 $0.30 $1.80
Grok Imagine Lite, 480p $0.02 $0.20 $1.20
Grok Imagine 1.5, 1080p $0.25 $2.50 $15.00
Veo 3.1, 1080p with audio $0.40 $4.00 $24.00

One minute of waiting for ten seconds of video

Artificial Analysis reports a median generation time of 60.5 seconds for a 10-second 1080p Lite clip. Among the 12 models included in its speed comparison, Lite occupies the quality-speed frontier.

Kling 3.0 1080p Pro scores slightly higher on quality and takes 94 seconds to produce a 5-second clip. Vidu Q3 Turbo’s listed 5-second 720p configuration finishes about nine seconds sooner than Lite’s 10-second 1080p configuration and receives a substantially lower quality score. The differing durations and resolutions limit direct throughput comparisons, so these figures describe the tested configurations.

A roughly one-minute turnaround supports iterative prompt development, storyboarding, and variant generation without scheduling long batch runs. Actual API latency can still vary with provider queues, retries, region, and request settings.

Narrative and lighting lead the scores

The AA-Video-T2V v2.0 benchmark evaluates ten capabilities and ten use cases. Lite comes closest to the quality frontier in three capability groups:

  • Multi-Scene & Narrative: Multi-event stories, shot transitions, montage, and identity consistency across shots.
  • Lighting & Materials: Exposure, shadows, color temperature, cloth, hair, and fur.
  • Text Rendering: Short and long text, two-dimensional layout, and text placed on deforming surfaces.

Its largest deficits appear in Dialogue & Lip Sync and Human Anatomy. Compared with the full Grok Imagine Video 1.5, Lite matches the Physics score and trails across the other nine capability groups. The smallest gap between the two models appears in Multi-Scene & Narrative.

Use-case scores favor Architecture & Real Estate, Consumer content, and Productivity & Knowledge Work. Those categories cover exterior fly-throughs, interior walkthroughs, virtual staging, selfies, personal footage, b-roll, explainers, educational clips, and animated data visualization.

Live-Action Film and the benchmark’s Frontier category produce Lite’s largest gaps. Dialogue-heavy scenes, close shots of hands or faces, and demanding cinematic sequences therefore warrant prompt-specific evaluation before production use.

From prompt to MP4

Lite is available through the xAI API, fal, and Vercel AI Gateway. Fal lists the same base prices. Vercel exposes the model through the AI SDK’s experimental video helper.

With an AI Gateway key configured in the AI_GATEWAY_API_KEY environment variable, a minimal Node.js request looks like this:

javascript
import { experimental_generateVideo as generateVideo } from 'ai';
import { writeFile } from 'node:fs/promises';

const result = await generateVideo({
  model: 'xai/grok-imagine-video-1.5-lite',
  prompt:
    'A drone shot of a coastal cliff with crashing waves at golden hour.',
  duration: 5,
  providerOptions: {
    xai: {
      pollTimeoutMs: 600000,
    },
  },
});

const video = result.videos[0];

if (video === undefined) {
  throw new Error('The provider returned no video.');
}

await writeFile('output.mp4', video.uint8Array);

The pollTimeoutMs value allows the SDK to poll the asynchronous job for up to ten minutes. The example leaves resolution and aspect ratio at the route’s defaults. Supported request fields can differ across xAI, fal, and Vercel, so applications should validate the selected route’s current options before exposing controls to users. The helper remains experimental, making an explicit AI SDK version useful for production deployments.

Provider catalogs also differ in how they expose the 1080p tier. OpenRouter identifies it as an upscaled 720p render, while some third-party listings expose only 480p and 720p for the Lite model ID. Applications that depend on fine 1080p detail should compare Lite’s upscaled output with the full model’s output using representative prompts.

A practical two-tier pipeline

A two-tier workflow can use Lite for prompt exploration and reserve the full model for selected prompts. Ten 10-second Lite drafts at 1080p cost $14 at the published rate. Generating one selected prompt again with the full model adds $2.50, bringing the total to $16.50. Ten full-model drafts would cost $25.

Using 720p for the draft stage reduces that example to $3 for ten drafts, plus $2.50 for one full-model 1080p generation. Each rerun remains a fresh generation, and exact repeatability depends on any seed controls exposed by the provider.

Lite’s benchmark profile aligns most closely with architectural walkthroughs, explainers, social clips, b-roll, and multi-scene drafts. The full model or a stronger specialist remains the safer budget allocation for synchronized dialogue, anatomy-sensitive close-ups, and demanding live-action sequences.

The release gives developers a lower-cost tier between basic video tools and frontier-priced generators. Its position above Veo 3.1 on this leaderboard adds price pressure, although the ranking reflects a specific prompt set, scoring method, and model snapshot. Production evaluations should also measure route availability, failed-job behavior, queue latency, audio quality, and performance on an application’s own prompts.

Trending
  • No trending articles

Comments

avatar

Next Reads