ElevenLabs' Scribe v2 Now Edits Transcripts Without a Separate AI Call
ElevenLabs adds natural-language instructions to Scribe v2 transcription, letting you reformat, redact, or annotate transcripts in one API call for a 30% surcharge.
- ElevenLabs launched Transcript Editing for Scribe v2 and Scribe v2 Realtime speech-to-text.
- Pass a natural-language
transcript_editinstruction, up to 2,000 characters, per request. - API returns both original transcript with word timestamps and an
edited_transcriptfield. - Supports date/time normalization, abbreviation expansion, redaction, sentiment tagging, and full reformatting.
- Costs 30% on top of base transcription, with a 10-second minimum billed per request.
- Cannot combine with entity detection, entity redaction, or multi-channel transcription requests.
ElevenLabs has added instruction-based post-processing to its speech-to-text API. The experimental Transcript Editing feature accepts a plain-English editing instruction with a Scribe v2 or Scribe v2 Realtime request, then returns the original transcript alongside a rewritten version.
Developers can use the feature to normalize dates, expand abbreviations, mask sensitive language, add labels, or restructure text without making a separate large language model call. That removes a network round trip and reduces the prompt, retry, and billing logic required for a separate post-processing service.
One request, two outputs
A batch request includes a transcript_edit parameter of up to 2,000 characters. Scribe first transcribes the audio, then applies the instruction across the completed transcript wherever it is relevant.
| Output | Contents | Typical use |
|---|---|---|
text |
Original transcript | Search, auditing, and source retention |
| Word timestamps | Original timing data | Captions, alignment, and indexing |
edited_transcript |
Instruction-shaped text and status information | Display, export, or downstream processing |
The editing pass leaves the original text and word-level timestamps unchanged. Applications that require exact alignment should continue using those original fields because rewritten text can differ in wording and structure.
Edits suited to transcript cleanup
ElevenLabs documents several editing patterns that fit the feature’s transcript-focused scope:
- Normalize spoken dates as
YYYY-MM-DDor convert times to 24-hour notation. - Expand abbreviations on first use, such as changing
ETAtoestimated time of arrival (ETA). - Mask profane or sensitive words with a specified replacement pattern.
- Attach sentiment labels selected from a fixed list.
- Restructure a transcript as a bulleted list with one sentence per item.
One instruction can combine several operations, and the editor applies each applicable rule throughout the transcript rather than stopping after the first match.
A batch request in Python
The Python SDK exposes transcript editing as an additional argument on the standard conversion call:
transcription = elevenlabs.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
transcript_edit="Write all dates in ISO 8601 format (YYYY-MM-DD)",
)
print("Original:", transcription.text)
print("Edited:", transcription.edited_transcript)The response includes an edited_transcript object with a kind field. A successful edit contains the rewritten text. If editing fails after transcription succeeds, the API still returns the original transcript, allowing the application to preserve or process the unedited result.
Realtime editing waits for committed text
Scribe v2 Realtime accepts the instruction when the WebSocket connection opens. Each committed transcript is then followed by an edited_transcript event, as described in the realtime guide.
Partial transcripts are excluded from the editing pass. Processing only committed text reduces repeated work and prevents an interface from displaying rewrites of incomplete utterances.
The surcharge and hard limits
Transcript editing adds 30% to the base transcription cost, with billing subject to a minimum of 10 seconds of audio per request. Teams evaluating the feature should also account for these constraints:
- The edit begins after transcription, so total latency increases with transcript length.
- Instructions can use any language, although ElevenLabs says English performs best.
- The edited output remains in the source language.
- Spoken audio is treated as data rather than as editing instructions, limiting prompt-injection attempts embedded in the recording.
- A request cannot combine transcript editing with
entity_detection,entity_redaction, oruse_multi_channel.
Use cheaper controls first
Several deterministic transcription options cover common cases at lower cost and with more predictable behavior:
keytermprompting biases recognition toward names, product terms, and specialized vocabulary.no_verbatimremoves filler words and disfluencies.numbers_format, where supported, controls whether numbers appear as digits or words.
Transcript editing is better suited to transformations those controls cannot express, including schema-ready dates, fixed sentiment labels, profanity masking, and UI-specific formatting.
Where it fits in a pipeline
Applications that already send every transcript to a separate language model may gain little from moving the edit into Scribe. A dedicated post-processing stage provides control over model selection, prompt versioning, caching, validation, and intermediate output.
Applications with narrow, repeatable cleanup requirements can remove that extra orchestration step by using the integrated editor. The decision depends on whether the 30% surcharge costs less than operating a separate model call and whether Scribe’s editing behavior provides enough control for the required output.
Normalization, redaction, labeling, and structural reformatting fit the feature’s design. Summarization, inference, and multi-step reasoning remain stronger candidates for a dedicated downstream model.