Fish Audio's transcribe-1-pro Adds Speaker ID and Emotions for $0.36 an Hour

Fish Audio's transcribe-1-pro adds inline speaker markers and emotion cues like [laughter] to transcripts across 83 languages, at $0.36 per audio hour.

·
·
Read4 min
TypeNews
  • Fish Audio launched transcribe-1-pro, an ASR model with inline speaker labels and emotion cues.
  • Speakers appear as <|speaker:N|> markers, emotions as bracketed tags like [laughter].
  • Priced at $0.36 per audio hour, matching the standard transcribe-1 model.
  • Supports 83 languages with automatic language detection and word-level timestamps.
  • API-only for now, selected via the model: transcribe-1-pro HTTP header on POST /v1/asr.
  • Response is one annotated string, not a structured turns array, so parsing is on you.

Fish Audio’s Pro transcription adds speakers and vocal cues

Fish Audio has added speaker diarization, which identifies who spoke when, and vocal-event annotation to its speech-to-text API through transcribe-1-pro. The model inserts speaker changes, emotions, and sounds such as laughter into transcripts while returning automatic language detection and word-level timestamps. Meeting, support, podcast, and voice-agent systems can capture those signals in one recognition pass.

Both transcribe-1-pro and the standard transcribe-1 use POST /v1/asr. Existing clients select Pro through an HTTP header, leaving the endpoint and request structure unchanged.

One string carries every turn

The response’s text field uses inline <|speaker:N|> markers, where N is a numeric speaker identifier. Each marker applies until the next one appears, and the same identifier returns when that speaker resumes. Applications must map those identifiers to names when needed. Emotions and vocal events appear as bracketed annotations beside the relevant words.

css
{
  "text": "<|speaker:0|>Hello.<|speaker:1|>[happy]Nice to meet you.<|speaker:0|>Same here.",
  "duration": 6.4,
  "segments": [],
  "language_code": "en",
  "language": "English"
}

When timestamps are requested, the segments data contains spoken words with speaker, emotion, and event markers removed. Segments have no per-segment speaker_id, so captioning systems that need timed speaker labels must reconcile the annotated transcript with the timestamped text.

A small regular expression can convert the annotated string into turn objects for downstream processing:

python
import re

pattern = r"<\|speaker:(\d+)\|>(.*?)(?=<\|speaker:\d+\|>|$)"

for turn in re.finditer(pattern, result["text"], re.DOTALL):
    speaker_id = turn.group(1)
    text = turn.group(2).strip()
    print(f"Speaker {speaker_id}: {text}")

One header enables Pro

Existing integrations can activate the model by adding model: transcribe-1-pro to the request. Authentication, audio upload, and timestamp controls remain on the same API call:

nginx
curl --request POST https://api.fish.audio/v1/asr \
  --header "Authorization: Bearer $FISH_API_KEY" \
  --header "model: transcribe-1-pro" \
  --form audio=@speech.wav \
  --form ignore_timestamps=false

The Python SDK routes model selection through RequestOptions.additional_headers={"model": "transcribe-1-pro"} because the model is not exposed as a first-class method argument.

The endpoint accepts formats including WAV, MP3, and Opus through multipart form data or MessagePack request bodies. Each request processes one file. Applications that divide long recordings across requests must add cumulative timestamp offsets when assembling the results.

The bill stays at 36 cents

  • Rate: transcribe-1-pro costs $0.36 per audio hour, matching the standard model.
  • Billing unit: Fish Audio charges by the second of processed audio, independent of token count or request count.
  • Access: Pro is currently available through the API, with web-app support planned.
  • Languages: The model covers 83 languages and detects the spoken language automatically.
  • Language hints: The language parameter guides detection, which can still return a different language.

Why the extra tokens matter

A conventional conversation pipeline may run transcription, speaker diarization, and vocal-event classification as separate stages, then merge their outputs. Pro emits those signals together. Meeting summaries can attribute decisions, support analytics can associate reactions with specific turns, and voice-agent evaluations can use annotations such as [laughter], [surprised], or [frustrated] as additional scoring inputs.

The compact response shifts some work into application code. Developers must parse speaker turns, map numeric labels to known participants, and align annotated text with timestamped segments. The bracketed cues are model predictions and require validation against the accents, recording conditions, overlapping speech, and vocabulary found in the target workload.

Calls, podcasts, and agent loops

Fish Audio’s broader voice stack includes the S2.1 Pro text-to-speech model, for which the company reports time to first audio near 90 milliseconds, support for 83 languages, and cross-language voice identity. Pairing that generator with transcribe-1-pro gives applications an API for capturing conversational context before an LLM, agent, or synthesis stage processes it.

  • Meetings and calls: Attribute decisions, questions, and follow-up tasks to individual speakers.
  • Podcasts and interviews: Retain laughter, sighs, and delivery cues in searchable transcripts.
  • Voice-agent evaluation: Add speaker and reaction signals to conversation scoring.
  • Dubbing and localization: Carry delivery cues into translated scripts and regenerated audio.

At the listed rate, a 60-minute recording costs $0.36 to process. Teams currently combining transcription with separately billed diarization or emotion analysis can compare that single-pass cost against their existing pipeline. The main integration expense is adapting systems that expect speaker-aware segment objects to Fish Audio’s annotated-string format.

Trending
  • No trending articles

Comments

avatar

Next Reads