HiDream.ai's HiDream-O1-Video-1.0 Debuts at No. 6 Generating 1080p Videos With Synced Audio
HiDream's new omnimodal video model lands at #6 on the Artificial Analysis I2V with Audio leaderboard, generating 1080p clips with synchronized audio at $5.80 per minute.
- HiDream-O1-Video-1.0 debuts at #6 on the Artificial Analysis I2V with Audio leaderboard at Elo 1,175
- Generates 1080p clips of 5 to 20 seconds with natively synchronized audio, duration chosen by the model
- Priced at $5.80 per minute (about $0.10 per second) through the HiHarness API
- Public API only accepts one reference image plus optional text prompt, no audio or resolution controls
- Built on HiDream's Unified Transformer (UiT) foundation alongside image, world and embodied model families
- Available via HiHarness API and vivago R1 Studio with async task workflow
Beijing-based HiDream.ai has entered the top 10 of the Artificial Analysis leaderboard. HiDream-O1-Video-1.0 debuted at No. 6, behind Dreamina Seedance 2.0 720p and ahead of Wan 3.0. The nearby cluster includes MiniMax H3, Vidu Q4 Preview, and Gemini Omni Flash.
The top-six debut gives developers another competitive option for generating short videos with synchronized audio from a reference image. HiDream describes the model as natively omnimodal, with text, video, and audio processed together during generation. Its public API offers a narrower feature set than the launch materials describe, which limits the workflows developers can deploy today.
One pass for motion and sound
According to the launch materials, HiDream-O1-Video-1.0 generates 1080p videos lasting 5 to 20 seconds. The model can select a duration based on the action, event sequence, and pacing of the requested scene.
HiDream says the generation process accounts for gravity, inertia, collisions, deformation, materials, lighting, and spatial continuity. Its joint audiovisual architecture uses visual motion to inform sound timing while dialogue, effects, and text prompts shape the scene. Those claims still require testing across specific content types, especially scenes involving complex interactions or precise synchronization.
The API exposes one path
Developers using HiHarness must provide exactly one reference image as a public URL or Base64 payload. An optional text prompt can accompany the image and defaults to an empty value.
| Capability | Current API behavior |
|---|---|
| Reference image | Exactly one public URL or Base64 payload |
| Image size | 40 KB to 20 MB |
| Text prompt | Optional; defaults to empty |
| Duration | Selected by the model or fixed at 10 seconds with force_10s |
| Aspect ratio | Preserves the source image’s ratio |
| Resolution control | Unavailable |
| Reference video | Unsupported by the public endpoint |
| Audio control | Generated with the video; no public control field |
| Output | One video |
Launch materials list text, image, and video as possible inputs, while the current HiHarness documentation describes image-to-video generation only. Production designs should follow the documented request schema and verify output behavior before relying on broader launch claims.
Polling hides success in a subtask
HiHarness runs generation asynchronously, so an integration must submit a job, store its task_id, and poll for completion. The parent and subtask statuses carry different meanings:
- Submit the generation request.
- Read and persist the returned
task_id. - Poll the result endpoint until
result.statusequals1. - Treat that value only as confirmation that every subtask has stopped.
- Inspect each
sub_task_results[].task_statusto determine the outcome.
task_status = 3: generation failuretask_status = 4: safety-review failure
Production integrations should retain the raw subtask status and error payload, apply timeouts to polling, and avoid reporting success from the parent status alone.
The bill follows clip length
| List price | $5.80 per minute of 1080p video with audio |
|---|---|
| Approximate rate | $0.097 per second |
| API access | HiHarness |
| Studio access | vivago R1 Studio |
At the listed rate, linear cost estimates are about $0.48 for five seconds, $0.97 for 10 seconds, and $1.93 for 20 seconds. Model-selected duration introduces variable per-job cost, while force_10s provides a predictable clip length.
No. 6, with room to move
HiDream-O1-Video-1.0 has a leaderboard score of 1,175 ± 10 across 3,231 samples. Its confidence interval is the widest among the top six models, reflecting a smaller evidence base than some established competitors.
The wider interval leaves more room for the ranking to change as additional head-to-head votes arrive. Artificial Analysis captures evaluator preference through blind comparisons; physical accuracy, prompt adherence, latency, API reliability, and controllability require separate testing.
Four families share one foundation
HiDream places the video model in a portfolio built on UiT, or Unified Transformer, which the company describes as a common neural-network foundation:
- HiDream-O1-Image: image understanding and generation
- HiDream-O1-Video: video generation and temporal storytelling
- HiDream-O1-World: 3D environment simulation and real-time interaction
- HiDream-O1-Embodied: spatial reasoning, action planning, and feedback for embodied AI
Best fit: short image-led clips
| Workflow requirement | Fit |
|---|---|
| Single image to short video | Supported |
| Generated synchronized audio | Supported, without public controls |
| Fixed 10-second output | Supported through force_10s |
| Text-to-video | Absent from the current public endpoint |
| Multiple reference images | Unsupported |
| Video-to-video editing | Unsupported |
| Resolution selection | Unavailable |
| Audio editing or configuration | Unavailable |
Teams producing product shots, social clips, or storyboard animatics from a single still have the clearest use case. A useful evaluation should compare motion coherence, audiovisual timing, prompt adherence, latency, failure rates, and total cost against Seedance 2.0, MiniMax H3, and Wan 3.0. Projects requiring text-only generation, multiple references, video editing, or detailed output controls need another endpoint unless HiDream expands the public API.