Higgsfield's Seed Audio 1.0 Brings Voice Cloning and Dubbing Into Claude
Higgsfield's Seed Audio 1.0 brings voice cloning, TTS narration, and 18-language video dubbing to Claude via MCP — no API key required
- Seed Audio 1.0 is live on Higgsfield and via Claude MCP — covers TTS narration, voice cloning, and video dubbing in 18+ languages.
- Claude integration via MCP: connect at higgsfield.ai/mcp with one URL, no API key needed, using your existing Higgsfield credits.
- Multi-model audio stack: Seed Audio 1.0 (ByteDance) and ElevenLabs v3 power the engine, with MiniMax and Vibe Voice also available under one subscription.
- Voice cloning from a short audio sample supports 74+ languages and can be reused across narration, ads, dubbing, and multi-speaker dialogue.
- Full scene audio generation: Seed Audio 1.0 generates multi-speaker dialogue, ambience, background music, and foley effects in a single pass.
- Free tier includes 150 credits/month; paid credits at ~$1 per 16 credits, with Claude Code one-command setup:
claude mcp add --transport http --scope user higgsfield https://mcp.higgsfield.ai/mcp
Higgsfield just shipped Seed Audio 1.0, a new audio model that handles three distinct tasks in one place: text-to-speech narration, voice cloning from a short sample, and full video dubbing across 18 languages. It is available directly on the Higgsfield platform and , the more interesting part for developers , through Claude via the Higgsfield MCP server.
What Seed Audio 1.0 actually does
Seed Audio 1.0 is ByteDance Seed's all-in-one audio generation model for creating complete sound scenes. It accepts text, image, or audio context to guide multi-speaker dialogue, emotional delivery, native accents, ambience, background music, and foley-style effects. That last part matters: this is not just a TTS engine. It is positioned for complete audio scenes , multi-character dialogue, emotion, tone, accents, ambience beds, BGM, and foley in a single creative pass.
The three core modes are:
- Text-to-speech (Voiceover): Paste a script, pick a voice preset or your own clone, and get a narrated audio file.
- Voice swap: Replace the voice in an existing video with any preset or cloned voice.
- Video dubbing: Translate a video into a target language with automatic lip-sync to the new audio track.
Higgsfield brings ElevenLabs, MiniMax, Seed Speech, and Vibe Voice into one subscription, with multilingual cloning, dubbing, video sync, and a multi-voice library. Seed Audio 1.0 is the new addition to that roster, sitting alongside ElevenLabs v3 as a primary engine. Claude can now generate audio via Higgsfield MCP , voiceovers, voice cloning, and dubbing in 50+ languages, powered by Seed Audio 1.0 and ElevenLabs v3, all inside Claude.
The MCP angle is the real story
Higgsfield's MCP server has been live since late April, but audio was missing from it until now. Adding Seed Audio 1.0 completes the loop: you can now go from a written brief to a finished, dubbed video entirely inside a Claude conversation. Once the two are connected, Claude can reach every model and feature on Higgsfield from inside any conversation , you describe the shot you want and Claude takes care of choosing the right model, configuring the parameters, firing the generation, and bringing the finished clip back to chat.
Setup is minimal. The server is hosted at https://mcp.higgsfield.ai/mcp, speaks the Model Context Protocol over HTTP, authenticates through your existing Higgsfield account, and requires no API key to provision or developer app to register. For Claude Code users, the install is a single terminal command:
claude mcp add --transport http --scope user higgsfield https://mcp.higgsfield.ai/mcp
Higgsfield MCP works with Claude (web, Cowork, and Claude Code), OpenClaw, Hermes Agent, and NemoClaw , any agent or client that supports MCP can connect.
Voice cloning in practice
AI Voice Cloning learns a voice from a short sample and reads any script in that voice, cloning your tone for narration, ads, audiobooks, and dubbing across 74+ languages, so you keep one consistent voice across every project without a studio. The workflow is straightforward:
- Record or upload a clean audio sample (WAV or MP3, up to 2 minutes)
- Confirm consent and submit , the model learns tone, pitch, and style
- Paste any script, pick a language, tune tone and pace, generate
The voice cloning workspace brings every leading audio model into one place , you can switch between models per clone, compare results side by side, and pick the best take, all under one subscription with new releases added as they launch.
What it is good at , and where it falls short
Seed Audio 1.0 is designed for complete audio scenes: multi-character dialogue, emotion, tone, accents, ambience beds, BGM, and foley in a single creative pass. It composes multiple sound layers at once instead of stitching voice, music, and ambience in separate tools, and guides tone, emotional delivery, dialect, and native-sounding accents while keeping recurring voices recognizable across contexts.
The dubbing feature is where most creators will find the most immediate value. The Translate feature allows educational, corporate, and social media content to scale globally in minutes , instantly converting an English video into Mandarin, Hindi, French, or Japanese, complete with native lip-sync, so businesses can multiply their viewership without multiplying their production workload. That said, for complex dialogue-heavy content, providing a reference audio clip gives the best results. The model also requires a clearly visible face in the video for the lip-sync to look professional.
Practical use cases
The combination of audio generation and the existing video toolkit in Higgsfield MCP opens up some concrete workflows:
- Localized product videos: Generate a product launch video, then dub it into 18 languages in the same session without leaving Claude.
- Faceless YouTube channels: Write a script, clone a narrator voice, generate visuals, and combine them , no recording studio needed.
- Ad variant testing: Swap voices across different audience-targeted versions of the same video to A/B test tone and delivery.
- E-learning localization: Turn heavy written manuals into engaging video presentations using Voiceover, then use the Translate feature to localize training materials for global branches in languages like Russian, Portuguese, Turkish, and more.
- Podcast and multi-speaker dialogue: Vibe Voice handles cloned voices in multi-speaker dialog and conversational reads, ideal for podcasts and dialogue.
Pricing and access
Seed Audio 1.0 is available now on the Higgsfield platform and through the MCP server. Higgsfield tools use the same credit system as the Higgsfield platform, with each generation costing credits based on the model and resolution. The free tier gives 150 credits per month. New Higgsfield accounts ship with free credits, so you can run your first generations without committing to a paid plan. Beyond that, prepaid credits run at approximately $1 for 16 credits.
For teams already using the Higgsfield MCP for image and video generation, audio is now just another tool call in the same conversation , no new subscriptions, no new authentication, no context switching. That frictionless integration into an agentic workflow is what makes this update worth paying attention to.