ElevenLabs' Scribe v2 Now Edits Transcripts Without a Separate AI Call

ElevenLabs adds natural-language instructions to Scribe v2 transcription, letting you reformat, redact, or annotate transcripts in one API call for a 30% surcharge.

·
·
Read4 min
TypeNews
TopicAudio · Api
  • ElevenLabs launched Transcript Editing for Scribe v2 and Scribe v2 Realtime speech-to-text.
  • Pass a natural-language transcript_edit instruction, up to 2,000 characters, per request.
  • API returns both original transcript with word timestamps and an edited_transcript field.
  • Supports date/time normalization, abbreviation expansion, redaction, sentiment tagging, and full reformatting.
  • Costs 30% on top of base transcription, with a 10-second minimum billed per request.
  • Cannot combine with entity detection, entity redaction, or multi-channel transcription requests.

ElevenLabs has added instruction-based post-processing to its speech-to-text API. The experimental Transcript Editing feature accepts a plain-English editing instruction with a Scribe v2 or Scribe v2 Realtime request, then returns the original transcript alongside a rewritten version.

Developers can use the feature to normalize dates, expand abbreviations, mask sensitive language, add labels, or restructure text without making a separate large language model call. That removes a network round trip and reduces the prompt, retry, and billing logic required for a separate post-processing service.

One request, two outputs

A batch request includes a transcript_edit parameter of up to 2,000 characters. Scribe first transcribes the audio, then applies the instruction across the completed transcript wherever it is relevant.

Output Contents Typical use
text Original transcript Search, auditing, and source retention
Word timestamps Original timing data Captions, alignment, and indexing
edited_transcript Instruction-shaped text and status information Display, export, or downstream processing

The editing pass leaves the original text and word-level timestamps unchanged. Applications that require exact alignment should continue using those original fields because rewritten text can differ in wording and structure.

Edits suited to transcript cleanup

ElevenLabs documents several editing patterns that fit the feature’s transcript-focused scope:

  • Normalize spoken dates as YYYY-MM-DD or convert times to 24-hour notation.
  • Expand abbreviations on first use, such as changing ETA to estimated time of arrival (ETA).
  • Mask profane or sensitive words with a specified replacement pattern.
  • Attach sentiment labels selected from a fixed list.
  • Restructure a transcript as a bulleted list with one sentence per item.

One instruction can combine several operations, and the editor applies each applicable rule throughout the transcript rather than stopping after the first match.

A batch request in Python

The Python SDK exposes transcript editing as an additional argument on the standard conversion call:

vim
transcription = elevenlabs.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    transcript_edit="Write all dates in ISO 8601 format (YYYY-MM-DD)",
)

print("Original:", transcription.text)
print("Edited:", transcription.edited_transcript)

The response includes an edited_transcript object with a kind field. A successful edit contains the rewritten text. If editing fails after transcription succeeds, the API still returns the original transcript, allowing the application to preserve or process the unedited result.

Realtime editing waits for committed text

Scribe v2 Realtime accepts the instruction when the WebSocket connection opens. Each committed transcript is then followed by an edited_transcript event, as described in the realtime guide.

Partial transcripts are excluded from the editing pass. Processing only committed text reduces repeated work and prevents an interface from displaying rewrites of incomplete utterances.

The surcharge and hard limits

Transcript editing adds 30% to the base transcription cost, with billing subject to a minimum of 10 seconds of audio per request. Teams evaluating the feature should also account for these constraints:

  • The edit begins after transcription, so total latency increases with transcript length.
  • Instructions can use any language, although ElevenLabs says English performs best.
  • The edited output remains in the source language.
  • Spoken audio is treated as data rather than as editing instructions, limiting prompt-injection attempts embedded in the recording.
  • A request cannot combine transcript editing with entity_detection, entity_redaction, or use_multi_channel.

Use cheaper controls first

Several deterministic transcription options cover common cases at lower cost and with more predictable behavior:

  • keyterm prompting biases recognition toward names, product terms, and specialized vocabulary.
  • no_verbatim removes filler words and disfluencies.
  • numbers_format, where supported, controls whether numbers appear as digits or words.

Transcript editing is better suited to transformations those controls cannot express, including schema-ready dates, fixed sentiment labels, profanity masking, and UI-specific formatting.

Where it fits in a pipeline

Applications that already send every transcript to a separate language model may gain little from moving the edit into Scribe. A dedicated post-processing stage provides control over model selection, prompt versioning, caching, validation, and intermediate output.

Applications with narrow, repeatable cleanup requirements can remove that extra orchestration step by using the integrated editor. The decision depends on whether the 30% surcharge costs less than operating a separate model call and whether Scribe’s editing behavior provides enough control for the required output.

Normalization, redaction, labeling, and structural reformatting fit the feature’s design. Summarization, inference, and multi-step reasoning remain stronger candidates for a dedicated downstream model.

Trending
  • No trending articles

Comments

avatar

Next Reads