Shengshu's Vidu Q4 Preview Debuts at No. 3 for AI Video
Shengshu's new flagship jumps 16 spots on the Artificial Analysis image-to-video leaderboard while costing 25 percent less per minute than Vidu Q3 Pro.
- Vidu Q4 Preview debuts at #3 on AA-Video-I2V v1.0 with 1,179 Elo, up from #19.
- Generates 3 to 16 second clips at up to 4K with native synchronized audio.
- Accepts up to 15 image references and 3 audio references for consistent characters.
- Priced at $7.20 per minute at 1080p, matching Q3 Pro at better quality.
- Supports image-to-video and reference-to-video only, no text-to-video yet.
- Available now via Vidu web and API, 30% discount through November.
Vidu Q4 Preview debuts at No. 3 for image-to-video
Shengshu Technology has released Vidu Q4 Preview, an early version of its next flagship video model. It entered Artificial Analysis’s AA-Video-I2V v1.0 leaderboard at No. 3 with an Elo score of 1,179. Vidu Q3 Pro, the company’s previous highest-ranked model, sits at No. 19 with 1,056, a difference of 16 places and 123 points.
Artificial Analysis derives Elo scores from pairwise human preferences between generated clips. The score measures relative performance within that test, while factors such as latency, API reliability, long-sequence consistency and workload-specific cost require separate evaluation. Leaderboard positions can also move as more votes arrive.
References drive the preview
Q4 Preview currently supports image-to-video and reference-to-video generation through the Vidu platform. Prompt-only text-to-video generation remains unavailable in this release.
| Capability | Q4 Preview specification |
|---|---|
| Clip length | Up to 16 seconds |
| Frame rate | 24 frames per second |
| Resolution | 540p, 720p, 1080p, 2K and 4K |
| Audio | Simultaneous video and audio generation |
| Image references | Up to 15 |
| Audio references | Up to three |
| Color depth | 10-bit output at 2K and 4K |
| Aspect ratio | Preserves the source image’s ratio |
The reference controls can anchor a character’s face, wardrobe, props, environment and featured products within one generation. Audio references provide a similar control for voices, helping maintain the same speaker across separate shots.
Faces, motion and camera control
Shengshu says Q4 Preview better coordinates facial expressions, body movement, emotion, speech and voice. Those changes target common video-generation failures, including drifting faces, poor lip synchronization, abrupt body movement and inconsistent performances between frames.
The model also aims to track fast action more coherently during fights and chases. Shengshu claims that explosions, fireworks and particle effects blend more naturally with their surroundings instead of appearing detached from the scene.
A tester quoted in the company’s launch announcement reported strong Chinese text rendering and less need for post-production effects. That observation comes from an early user rather than a controlled text-rendering benchmark.
The rate card needs decoding
Published prices describe several configurations and do not map cleanly onto one another. Artificial Analysis lists Q4 Preview at $7.20 per minute for its 1080p test configuration with audio, compared with $9.60 for Q3 Pro.
| Price reference | Per second | Per minute |
|---|---|---|
| Artificial Analysis, 1080p with audio | $0.12 | $7.20 |
| Company-advertised minimum | $0.014 | $0.84 |
| Company-listed 540p rate | $0.045 | $2.70 |
| Company-listed 4K rate | $0.39 | $23.40 |
The advertised minimum falls below the listed 540p rate, which indicates a different configuration or promotional billing condition. Shengshu’s public material does not fully specify how audio, duration, reference mode and resolution affect that minimum. The launch announcement also offers a 30% discount on image-to-video and reference-to-video generation through November 30, 2026, without clarifying whether every quoted rate includes the promotion.
Shengshu claims Q4 Preview can produce up to five times more output for the same budget under comparable specifications and billing conditions. The comparison uses an unreleased, full-price Q4 baseline, so the cost-efficiency claim cannot yet be independently verified.
A narrow lead pack
MiniMax H3 Max currently leads the I2V leaderboard with an Elo score of 1,195, followed by MiniMax H3 and Vidu Q4 Preview. Q4 trails the leader by 16 points, a gap smaller than the leaderboard’s roughly 20-point confidence intervals. Its 123-point advantage over Q3 Pro is considerably larger.
Runway does not appear within the board’s top 28 entries, while Luma Ray 3.2 is absent from the displayed rankings. Their omission leaves the leaderboard without a direct comparison against those models.
Best fits and hard limits
Q4 Preview’s current controls suit several production workflows:
- Animating still images and product shots into short clips with generated audio
- Maintaining characters, wardrobes and props across multiple cuts through reference images
- Producing 2K or 4K deliverables without a separate upscaling stage
- Testing Chinese titles and signs in scenes where text accuracy affects post-production work
- Generating short-form advertisements, action sequences and character-driven scenes
Prompt-only workflows require an image from another source before generation can begin. The 16-second limit also requires longer sequences to be assembled from multiple clips, with continuity managed across each cut.
Shengshu’s public materials omit several details needed for production integration, including endpoint names, SDK support, accepted codecs, file-size limits, deterministic controls, rate limits, queue latency and data-retention terms. Developers should verify those constraints, along with account-specific pricing and preview access, before building Q4 Preview into an automated pipeline.