API Docs

WebSocket API

Note: This is a consolidated document. For detailed specifications, refer to the individual documents under reference/websocket/.

Note: The URL used in this document (vas-poc.vurbo.ai) is the planned deployment address. A separate notice will be issued after the official launch.


Table of Contents

  1. Connection Info
  2. Authentication
  3. Message Format
  4. Health - Heartbeat Service
  5. Voice Translation - start
  6. Voice Translation - config
  7. Voice Translation - audio
  8. Voice Translation - pause
  9. Voice Translation - resume
  10. Voice Translation - stop
  11. Voice Translation - retranslate
  12. Voice Translation - switch_language
  13. Voice Translation - set_name
  14. Voice Translation - rename_speaker
  15. Voice Translation - reassign_speaker
  16. Voice Translation - merge_speakers
  17. Voice Translation - tts_play
  18. Voice Translation - tts_stop
  19. Voice Translation - tts_mode
  20. Voice Translation - set_tts
  21. Voice Translation - start_speaking
  22. Voice Translation - stop_speaking
  23. Voice Translation - switch_conversation_mode
  24. Voice Translation - set_speaker_language
  25. Voice Translation - set_speaking_speed
  26. Voice Translation - add_channel
  27. Voice Translation - remove_channel
  28. Voice Translation - set_channel_language
  29. Voice Translation - broadcast_go_live
  30. Voice Translation - broadcast_announcement
  31. Voice Translation - set_standby_message
  32. Response Events

Connection Info

ItemValue
Endpointwss://vas-poc.vurbo.ai/ws
ProtocolWebSocket
Data FormatJSON
Auth MethodTicket (see below)

Authentication

The VAS WebSocket uses a Ticket mechanism for authentication, passing a one-time Ticket via Sec-WebSocket-Protocol. For details, refer to Authentication.

Step 1: Obtain a Ticket

Exchange your API Key for a one-time Ticket via the REST API:

POST /api/v1/auth/ticket
X-API-Key: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

Response:

{
  "ticket": "aBcDeFgHiJkLmNoPqRsTuVwXyZ012345",
  "expires_in": 60
}
FieldTypeDescription
ticketstringOne-time Ticket (32 chars)
expires_inintValidity period (seconds)

Step 2: Connect to the WebSocket using the Ticket

Place the Ticket into Sec-WebSocket-Protocol in the format ticket.{TICKET_VALUE}:

// Native browser support
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);

ws.onopen = () => {
  console.log('Connected! Protocol:', ws.protocol);
  // Start using the WebSocket...
};

ws.onerror = (error) => {
  console.error('Connection failed:', error);
};

Node.js example:

const WebSocket = require('ws');

const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);

Ticket Characteristics

CharacteristicDescription
Validity period60 seconds
Usage countCan be used only once (deleted immediately after)
SecurityThe API Key is never exposed in the WebSocket connection
Replay protectionUses an atomic operation to guarantee single use

Ticket Error Codes

Error CodeHTTP StatusDescription
ticket_invalid401Ticket invalid or expired
ticket_expired401Ticket expired
ticket_already_used401Ticket already used
ticket_validation_failed500Ticket validation failed

For the full API specification, refer to Auth Ticket API.


Message Format

All messages use a unified nested structure:

{
  "type": "service type",
  "data": { ... }
}

Service Types

typeDescription
healthHeartbeat mechanism
voice-translationVoice translation service
errorError message

Error Message Format

When an error occurs, the server returns a message with type: "error":

{
  "type": "error",
  "data": {
    "error_code": "auth_invalid_api_key",
    "severity": "fatal",
    "message": "Invalid API key",
    "context": "auth",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-01-15T10:30:45.123Z"
  }
}

Sentence-level errors (such as a translation failure for one language of a sentence) additionally carry sid and details:

{
  "type": "error",
  "data": {
    "error_code": "llm_content_filtered",
    "severity": "warning",
    "message": "Content filtered",
    "context": "translation",
    "sid": 5,
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-01-15T10:30:45.123Z",
    "details": {
      "provider": "llm_service",
      "translation_language": "ja-JP"
    }
  }
}

Session-level translation service errors (escalated after consecutive failures reach a threshold) do not carry sid. The frontend should display a global notice but does not need to disconnect:

{
  "type": "error",
  "data": {
    "error_code": "translation_service_unavailable",
    "severity": "error",
    "message": "Translation service unavailable",
    "context": "translation",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-01-15T10:30:45.123Z",
    "details": {
      "provider": "llm_service",
      "last_error_code": "llm_provider_error",
      "fail_count": 5
    }
  }
}

For the full trigger rules (consecutive failure threshold, error code classification), refer to the translation_service_unavailable section in Error Code Reference.

Plan Limit Errors (unlimited plans, v1.9.0)

API Keys on an "unlimited plan" may receive the following error messages (not applicable to credit-based keys). Use GET /api/v1/me/plan to look up the plan contents, current usage, and when a restriction lifts:

Error CodeseverityDescriptionClient Handling
plan_feature_not_allowedfatalThe plan does not include the feature in useUse features included in your plan or upgrade the plan; query GET /api/v1/me/plan for the plan contents
concurrency_limit_reachederrorThis API Key has reached its concurrent recording limitThe connection is not closed; start again after another recording ends
daily_limit_disconnecterrorThe plan's usage threshold was reached; the current recording was stoppedYou may start a new recording immediately (when restarting on the same connection, session_started arrives after the previous recording finishes processing)
daily_limit_reachedfatalUsage has reached the plan's limitAvailable again after the plan's reset (daily limits reset the next day)

plan_feature_not_allowed has two occurrence points:

  1. start rejected: when the plan does not include a feature enabled in the request, the start is rejected outright; the connection is not closed — adjust the parameters and start again.
  2. Detected while recording: for example, a feature not in the plan is turned on mid-session; once the check before the next minute detects it, the current recording is stopped.

For real-time recording, usage limits are checked before each minute begins; once a limit is reached, the next minute does not start and is not counted toward usage. When the same API Key runs several recordings at once, after the periodic-stop threshold is reached, each recording stops before its own next minute begins.

In addition, details.max of too_many_languages may come from the plan's cap on simultaneously recognized transcription languages, in addition to the system-wide limit (10 transcription languages); details carries max and received. See Error Code Reference — Plan and Usage Limit Errors.

Single-Message Error (Message Failed, Connection Kept)

When the server encounters an unexpected internal error while handling a single WebSocket message (such as set_name, switch_language, tts_play, etc.), it returns internal_error. This error indicates only that the specific message failed to process; the connection is not terminated. The frontend should keep the connection open and may retry the operation:

{
  "type": "error",
  "data": {
    "error_code": "internal_error",
    "severity": "error",
    "message": "Internal server error",
    "context": "general",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-05-08T10:30:45.123Z",
    "details": {
      "message_type": "voice-translation",
      "action": "set_name"
    }
  }
}
details Fields
FieldTypeDescription
message_typestringService type: voice-translation / health
actionstring(Optional) The specific operation that failed, such as set_name, switch_language, tts_play, tts_mode, retranslate, config, speaker.rename, etc. This field is absent when the message payload has no action field (such as a plain init message).
What the Frontend Should Do
  1. Keep the WebSocket connection open: Do not call ws.close(), navigate away, or return to a history page because of this error. The recording is still in progress.
  2. Decide on follow-up handling based on details.action:
    ScenarioRecommended Action
    Idempotent operations such as set_name / switch_language / tts_mode / configSimply resend the same message. These operations use a "last write wins" approach, so retrying has no side effects.
    tts_play / tts_stop / retranslateUsually safe to retry directly. If the user is waiting for TTS playback, consider showing a transient toast indicating the retry is in progress.
    speaker.rename / speaker.mergeBefore retrying, use the REST API (speakers) to confirm the current DB state and avoid duplicate operations (for example, the rename already succeeded and only the response frame failed).
    details.action is absentThe error occurred after the message payload was parsed, so the system cannot infer the specific operation. The frontend can infer it from "the most recent message the user sent," or display a generic error message such as "Operation failed, please retry."
  3. User experience: Show a transient toast / inline error. Do not interrupt the user flow with a modal or a redirect.
  4. Telemetry / reporting: Report request_id + details to your frontend error tracking (Sentry, Datadog, etc.) to make it easier to correlate with backend logs during troubleshooting.
What Will Not Happen (Guarantees)
  • The recording will not be interrupted: segment_uploaded, result, origin, and other messages keep arriving.
  • The connection will not be actively closed by the server.
  • The session state will not be reset (session_id stays the same).
  • State already written to the DB will not be rolled back (for example, if set_name was written to the DB successfully and only the response frame failed, the name still takes effect).
Client Handling Example
ws.onmessage = (event) => {
  const msg = JSON.parse(event.data);
  if (msg.type !== 'error') {
    handleNormalMessage(msg);
    return;
  }

  const { error_code, severity, request_id, details } = msg.data;

  // Single-message failure: keep the connection, decide whether to retry based on action
  if (error_code === 'internal_error') {
    console.warn('[ws] message handler failed (connection kept)', {
      request_id,
      message_type: details?.message_type,
      action: details?.action,
    });
    showTransientToast(`Failed to process "${details?.action ?? 'operation'}", please retry`);
    // Note: do not call ws.close() and do not navigate away from the current page
    return;
  }

  // Handle other errors with your existing logic (only fatal errors require disconnecting)
  handleErrorBySeverity(severity, msg.data);
};
FieldTypeDescription
error_codestringError code (for programmatic handling)
severitystringSeverity: fatal / error / warning
messagestringHuman-readable error message
contextstringError source category
sidintOptional. The sentence number for sentence-level errors (such as a translation failure); absent for non-sentence-level errors
request_idstringRequest tracking ID
timestampstringTime the error occurred (ISO 8601)
detailsobjectOptional. Error context; common keys: provider, translation_language, source_lang, etc.

For the full list of error codes, refer to Error Code Reference.


Health (Heartbeat Service)

Description

Used to confirm that the WebSocket connection is healthy. We recommend sending a ping every 30 seconds; if no pong is received, treat the connection as dropped and reconnect. While the previous recording is still being processed after it ends (for example, while its summary is being generated), pong can be delayed by a few seconds to a few tens of seconds; allow enough waiting time and do not treat the connection as dropped just because a pong has not arrived yet.

Use Cases

  • Maintaining a long-lived connection
  • Detecting connection status
  • Preventing connection timeouts

Request - Ping

{
  "type": "health",
  "data": {
    "action": "ping"
  }
}

Response - Pong

{
  "type": "health",
  "data": {
    "action": "pong"
  }
}

Voice Translation - start (Start Voice Translation)

Description

Starts a new voice translation session and begins processing audio according to the configured parameters.

Use Cases

  • Starting a meeting record
  • Starting real-time translation
  • Starting a voice memo

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value start
transcription_languagesstring[]YesSpeech recognition languages (up to 10)
translation_languagesstring[]NoTranslation target languages; multiple allowed (up to 12; empty = no translation). As of v1.6.7, transcribe and broadcast translate all specified languages in real time, with one result event per language (same sid, single language key inside translations) — clients must accumulate by language code instead of overwriting. Two-way translation (conversation) does not apply: the server overwrites this field with the counterpart language, always a single language. record no longer supports translation as of v1.7.0 and returns 400 record_translation_not_allowed. See the WebSocket reference
realtime_translationbooleanNoReal-time translation mode (default false). true: translates word-by-word while the sentence is still being recognized (interim); false: translates only when the sentence is finalized. This flag also governs multi-language translation. Always treated as true for broadcast; has no effect for conversation, which translates finalized sentences only
recognition_modestringNoRecognition mode: single (single speaker, default), multi_speaker (multiple speakers), multi_channel (multi-channel, v1.10.0: one recording session takes input from multiple physical microphones, with speaker identity determined by the channel — see "Multi-Channel Mode Description" below). Under multi_speaker, transcription_languages must contain exactly 1 language; otherwise the server returns a diarization_multilang_conflict error and refuses to start (type=conversation is exempt: two-way translation forces single-speaker mode and has been exempt from this check since v1.7.2).
typestringYesRecording type: transcribe, conversation, record, broadcast
audio_formatstringNoAudio format: pcm (default), webm
summary_templatestringConditionalSummary template. Required for transcribe when summary_mode=builtin; forbidden when summary_mode=custom; optional for conversation/broadcast.
optionsobjectNoSpeech recognition options
tts_enabledbooleanNoWhether to enable TTS speech synthesis (default false)
tts_languagestringNoTTS output language (must be in translation_languages)
tts_voicestringNoTTS voice name (such as en-US-JennyNeural)
tts_modestringNoTTS playback mode: sync (synchronous, default), async (asynchronous). Only these two lowercase values are accepted, and an empty string is treated as not provided (sync); any other value (for example "Async") is rejected (invalid_parameter, details.field is tts_mode) and the recording does not start. Checked for every recording type, whether or not TTS is enabled
broadcast_tokenstringConditionalBroadcast token (required for the broadcast type, obtained from the REST API). Allowed only with the broadcast type: any other type that carries it is rejected (invalid_parameter, details.field is broadcast_token) and the recording does not start
active_languagestringNoInitial active language for two-way mode (default transcription_languages[0])
speakersarrayNoUser-to-language mapping for two-way mode (exactly 2 users when provided). When omitted, user 1 uses transcription_languages[0] and user 2 uses transcription_languages[1]
conversation_modestringNoTwo-way conversation mode: auto (auto-detect, default), manual (push-to-talk). Only these two lowercase values are accepted, and an empty string is treated as not provided (auto); any other value is rejected (invalid_parameter, details.field is conversation_mode) and the recording does not start. Also checked for recording types other than two-way
speaker_diarizationbooleanNoSpeaker diarization (forcibly ignored in two-way mode)
tts_configobjectNoMulti-language TTS settings (applies to both broadcast mode and two-way mode)
broadcast_phasestringNoInitial broadcast phase: standby, live (default). Only these two lowercase values are accepted, and an empty string is treated as not provided (live); any other value (for example "Live") is rejected (invalid_parameter, details.field is broadcast_phase) and the recording does not start
standby_messagestringNoThe message viewers see during the standby phase (default: "Getting ready, please wait...")
namestringNoInitial default recording name (max 60 chars after trimming leading and trailing whitespace; the system may still override it; if not provided, auto-generated such as Transcription #1). A longer name is rejected (invalid_parameter, details.field is name) and the recording does not start
summary_languagestringNoSummary output language (defaults to the recognition language when unspecified; in broadcast mode, read automatically from the channel settings). Up to 20 characters; a longer value is rejected (invalid_parameter, details.field is summary_language) and the recording does not start
summary_modestringNoSummary mode enum: builtin (default) / custom. Inferred as builtin when omitted.
summary_promptstringNoRequired in custom mode (a value with only whitespace counts as not provided); supplemental instructions in builtin mode. <= 3000 characters.
summary_prompt_slugstringNoRequired in custom mode (a value with only whitespace counts as not provided); forbidden in builtin mode. Your own identifier (<= 64 characters, Unicode, no control characters; passed through and stored in the backend record for historical lookup).
summary_plain_textbooleanNoRequest plain-text summary output (default false; when enabled, the backend performs Markdown post-processing).
channel_modestringConditionalMulti-channel sub-mode (required when recognition_mode=multi_channel): per_channel (each channel is recognized independently) or shared (channels take turns speaking and share one recognition stream). Other values return invalid_channel_mode
channelsarrayConditionalMulti-channel channel list (required when recognition_mode=multi_channel; includes the main speaker, conventionally channel_id: 1). See "Multi-Channel Mode Description" below
silenceTimeoutSecondsintegerNoHow many consecutive seconds without detected speech end the recording automatically: omitted or null uses the default (900 seconds); 0 means this session never ends for lack of speech; otherwise it must be an integer from 60 to 86400, and any other value is rejected. See Automatic End After a Long Silence below

options Sub-fields

options is the speech recognition options object; all fields are optional and use their respective defaults when omitted.

FieldTypeDefaultDescription
speaking_speedstringnormalSpeaking speed, which affects the silence threshold for sentence segmentation: very_slow / slow / normal / fast / very_fast. The default normal corresponds to an 800ms silence threshold (see speaking_speed levels for each level). Use a slower setting for slower speakers (longer threshold, avoids cutting on mid-sentence pauses); a faster setting segments sooner. Can be adjusted dynamically during recording via set_speaking_speed. Only these five lowercase values are accepted, and an empty string is treated as not provided (normal); any other value is rejected (invalid_parameter, details.field is options.speaking_speed) and the recording does not start
profanity_handlingstringmaskProfanity handling: mask (mask with ***) / remove (remove) / show (show original). Only these three lowercase values are accepted, and an empty string is treated as not provided (mask); any other value is rejected (invalid_parameter, details.field is options.profanity_handling) and the recording does not start

Note: These options apply to STT sentence segmentation; multi-speaker mode (multi_speaker) does not currently apply speaking_speed.

Recording Type Descriptions

typeDescriptionUse Cases
transcribeSpeech-to-textMeeting minutes, interview notes
conversationConversation recordTwo-way communication, customer service conversations
recordPlain recordingVoice memos, quick notes
broadcastBroadcast/liveLectures, talks, live content

Request Example (Basic)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "realtime_translation": false,
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_template": "meeting",
    "options": {
      "speaking_speed": "normal",
      "profanity_handling": "mask"
    }
  }
}

Request Example (Initial Default Name)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_template": "meeting",
    "name": "Product Planning Meeting"
  }
}

Recording Name Rules

ScenarioNamename_sourceSystem Override?
start with a name parameterInitial default namedefaultYes
start without a nameAuto-generated (such as Transcription #1, Broadcast #3)defaultYes
Set via set_nameThe name explicitly set by the useruserNo
Auto-generated by the system after the session endsA summary name generated from the transcript contentllm—

Note: The name in start is the initial default name; the system may still override it when the session ends. If you need a fixed name, use set_name.

Default name format (fixed English):

Recording TypeDefault Name Format
transcribeTranscription #N
conversationConversation #N
recordRecording #N
broadcastBroadcast #N

N is the sequential number for that user's recordings of the same type. Name priority: user > llm > default. Once the user sets a name, the system will not override it when the session ends.

Request Example (With TTS)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "realtime_translation": true,
    "type": "transcribe",
    "tts_enabled": true,
    "tts_language": "en-US",
    "tts_voice": "en-US-JennyNeural",
    "tts_mode": "sync"
  }
}

Request Example (Two-Way Mode - Auto-Detect)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "conversation",
    "transcription_languages": ["zh-TW", "en-US"],
    "active_language": "zh-TW",
    "audio_format": "pcm",
    "speakers": [
      { "id": 1, "language": "zh-TW" },
      { "id": 2, "language": "en-US" }
    ],
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
    }
  }
}

Request Example (Two-Way Mode - Manual Mode)

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "conversation",
    "transcription_languages": ["zh-TW", "en-US"],
    "conversation_mode": "manual",
    "audio_format": "pcm",
    "speakers": [
      { "id": 1, "language": "zh-TW" },
      { "id": 2, "language": "en-US" }
    ],
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-JennyNeural", "speaking_rate": 1.0 }
    }
  }
}

Request Example (Custom Summary Prompt - custom mode)

In mode=custom, your summary_prompt content completely replaces the built-in template rules, and the backend already adds prompt injection protection. The summary_prompt_slug is metadata for your own identification (stored in the backend record) and does not enter the prompt content.

If you want to keep the built-in template and add your own supplemental instructions afterward, use summary_mode=builtin + summary_template=<slug> + summary_prompt=<supplemental instructions> instead (in builtin mode, summary_prompt is treated as supplemental and appended after the built-in template).

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "transcription_languages": ["zh-TW"],
    "translation_languages": ["en-US"],
    "type": "transcribe",
    "audio_format": "pcm",
    "summary_language": "zh-TW",
    "summary_mode": "custom",
    "summary_prompt": "You are a meeting-minutes assistant. List every amount and committed date discussed in bullet points, and note the responsible person for each.",
    "summary_prompt_slug": "client_x_finance_v3",
    "summary_plain_text": false
  }
}

Important — How to Retrieve the Summary Result: In WebSocket mode, summaries are non-streaming by design; final_content is not pushed back via a WebSocket event (the summary_done event only signals completion and does not contain the content). The client must retrieve it afterward over HTTP:

  1. After receiving the summary_done event, call GET /api/v1/sse/history/transcribe/{taskId} to retrieve the summary (the init_summary event carries a top-level summary plain string + summary_mode / summary_template / summary_plain_text / summary_prompt_snapshot + the two content-filter fallback audit fields summary_fallback_level / summary_dropped_segments added in v1.5.5).
  2. Or query the summary_mode / summary_template / summary_prompt_slug fields of the transcript record via the REST API.

v1.5.5 Content-Filter Automatic Downgrade: If your prompt or transcript content triggers the LLM service's content filter, the system automatically downgrades (standard mode → neutral mode → segment-omission mode). The summary_fallback_level field of the summary_done event (value 2 or 3; omitted when standard mode succeeds directly) tells the client which path was actually taken, so the frontend can display hints such as "neutral mode in use" / "N segments omitted." See reference/websocket/events.md – summary_done and the V1.5.5 changelog.

Two-Way Mode Special Rules:

ItemDescription
transcription_languagesMust contain exactly 2 languages, and they cannot be the same.
translation_languagesNot required (automatically derived as the non-active language).
realtime_translationHas no effect in conversation mode: translations are sent only after a sentence is finalized, regardless of whether this field is true or false.
active_languageOptional, defaults to transcription_languages[0].
recognition_modeForced to single (ignores speaker_diarization).
tts_enabledDefaults to true; set to false to return text translations only.
tts_configOptional; sets the TTS voice for each of the two languages; leave empty to use the default voices automatically.
summary_templateOptional; when provided, a summary is automatically generated after stopping.
speakersOptional; specifies each user's language (exactly 2 users when provided). When omitted, users 1 and 2 map to the two transcription_languages in order.
conversation_modeOptional; auto (auto-detect, default) or manual (push-to-talk).

speakers Field Descriptions:

FieldTypeRequiredDescription
idintYesUser number (1 or 2)
languagestringYesThe user's language code (must be in transcription_languages)

conversation_mode Descriptions:

ModeDescription
auto (default)The system automatically detects the spoken language and segments sentences automatically.
manualThe user controls speaking periods via start_speaking / stop_speaking, during which the audio is merged into a single sentence.

Multi-Channel Mode Description (recognition_mode: "multi_channel")

A single recording session takes input from multiple physical microphones at the same time. Speaker identity is determined by the channel (not inferred by AI): each transcript sentence is attributed to the speaker of the channel_id it came from. There are two sub-modes:

  • per_channel: each microphone runs its own independent speech recognition; channels can speak at the same time and each can be bound to a different language.
  • shared: channels take turns speaking and share one recognition stream, and the whole session uses the same set of languages. Suited to situations where only one person speaks at a time; see Shared Mode below.

Scope and limits:

ItemDescription
Recording typesOnly transcribe and record; conversation returns invalid_parameter and broadcast returns multichannel_broadcast_not_allowed
Feature activationMulti-channel must be enabled before it can be used; in an environment where it is not enabled, start returns invalid_recognition_mode
channel_modeRequired; omitting it returns channel_mode_required. Allowed values are per_channel and shared; other values return invalid_channel_mode
channelsRequired (omitting it returns channels_required), 1–8 channels (including the main speaker, conventionally channel_id: 1); channel_id range 1–8, must be unique
Per-channel languageUnder per_channel, each channel must specify exactly 1 transcription language (each channel is bound to one language); the union of all channel languages (deduplicated) must exactly match the session-level transcription_languages, otherwise channel_language_mismatch is returned. Under shared, channels must not carry transcription_languages (returns channel_language_not_allowed); the whole session uses the session-level transcription_languages
audio_formatOnly pcm is supported (16kHz / 16-bit / mono / little-endian); other values return multichannel_requires_pcm
TTSNot supported in this first release: tts_enabled: true returns multichannel_tts_not_allowed
speaker_diarizationCannot be specified at the same time (multi-channel is itself a form of speaker diarization); providing both returns invalid_parameter
Plan limitsUnlimited plans must include the multi-channel feature; a plan may also cap the number of simultaneous recognition channels — exceeding the cap returns plan_feature_not_allowed on the spot at start (details.field="max_stt_streams"). shared is always counted as 1 recognition channel
Audio file length capThe audio saved for one multi-channel recording has a total cap, reached sooner the more channels are open (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration (duration_ms) stop at that point, while the transcript and credit charges continue as usual

channels field descriptions:

FieldTypeRequiredDescription
channel_idintYesChannel number, range 1–8, must be unique; includes the main speaker, conventionally 1 for the main speaker
speaker_namestringNoDisplay name of that channel's speaker (max 100 characters, no control characters, otherwise invalid_parameter is returned); when not provided, the speaker ID (channel_{N}) is displayed
transcription_languagesstring[]ConditionalTranscription language for that channel. Required under per_channel, exactly 1 (omitting it or providing more than one returns channel_language_required); must not be provided under shared (returns channel_language_not_allowed)

Multi-channel mode request example:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "per_channel",
    "transcription_languages": ["zh-TW", "en-US"],
    "translation_languages": ["ja-JP"],
    "audio_format": "pcm",
    "channels": [
      { "channel_id": 1, "speaker_name": "Manager Wang", "transcription_languages": ["zh-TW"] },
      { "channel_id": 2, "speaker_name": "Alex", "transcription_languages": ["en-US"] },
      { "channel_id": 3, "speaker_name": "Lee", "transcription_languages": ["zh-TW"] }
    ]
  }
}

Successful response (session_started, multi-channel):

The top level of data additionally carries channel_mode and channels[] (each entry with channel_id, speaker_name, transcription_languages, status; transcription_languages is not present in shared mode). The frontend should use this to confirm the server actually started in multi-channel mode:

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "per_channel",
    "channels": [
      { "channel_id": 1, "speaker_name": "Manager Wang", "transcription_languages": ["zh-TW"], "status": "preparing" },
      { "channel_id": 2, "speaker_name": "Alex", "transcription_languages": ["en-US"], "status": "preparing" },
      { "channel_id": 3, "speaker_name": "Lee", "transcription_languages": ["zh-TW"], "status": "preparing" }
    ],
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}

Multi-channel mode behavior summary:

  • Every audio frame must include channel_id (see the audio section); we recommend one frame every 100ms, and every channel should keep sending audio even while silent. A single silent channel does not end the recording; see Automatic End After a Long Silence
  • The origin of result events carries channel_id, speaker_id (format channel_{N}), and speaker_label (= speaker_name; after a rename_speaker, the new label); translations does not carry channel_id — use sid to match back to origin
  • During recording you can dynamically add or disable channels and change a channel's language via add_channel / remove_channel / set_channel_language; switch_language always returns multichannel_switch_language_not_allowed
  • rename_speaker works (a channel's speaker can be renamed before they say anything); reassign_speaker and merge_speakers do not apply to multi-channel (speaker identity is determined by the physical channel)
  • Channel status changes are reported via the channel_status event

Shared Mode (Taking Turns)

channel_mode: "shared": the microphones take turns speaking and share one recognition stream. The speaker is labeled by which channel that stretch of audio came in on, which is highly accurate when people take turns; billing is always counted as 1 channel (see Pricing).

Request example:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "transcribe",
    "recognition_mode": "multi_channel",
    "channel_mode": "shared",
    "transcription_languages": ["zh-TW", "en-US"],
    "audio_format": "pcm",
    "channels": [
      { "channel_id": 1, "speaker_name": "Host" },
      { "channel_id": 2, "speaker_name": "Guest A" },
      { "channel_id": 3, "speaker_name": "Guest B" }
    ]
  }
}

session_started returns channel_mode: "shared"; the entries in channels[] do not carry transcription_languages.

Client requirements:

ItemDescription
Send only one channel at a timeAt any moment, send audio only on the current speaker's channel. If several channels are sent at once, their audio is queued into the same recognition stream in the order received and the transcript becomes garbled (billing is not affected)
Keep sending silenceWhen nobody is speaking, keep sending silence on the current channel (one frame every 100ms recommended); do not stop sending
Audio formatpcm only (same as the common multi-channel rule)
LanguageThe whole session shares the session-level transcription_languages; channels cannot specify a language, and the language cannot be changed during the recording (set_channel_language returns channel_language_not_allowed)
The first channel cannot be removedThe first entry in channels[] carries the recognition for the whole session and cannot be removed with remove_channel (returns channel_remove_not_allowed); other channels can be removed

Channel status:

  • The status of the other channels follows the first channel: when the first channel turns ready, all channels turn ready together; when the first channel prepares recognition again (resume after pause, resume after disconnect, automatic reconnection), all channels return to preparing together. Each channel receives its own channel_status event, with the same reason as the first channel. After resuming from a disconnect, preparing is reported in the resume_ok snapshot rather than by a separate event.
  • A channel added with add_channel during the recording takes on the first channel's current status directly.
  • removed follows each channel's own status; a removed channel receives no further events.
  • When the first channel turns error, the other channels turn error as well (with the same reason); when the first channel recovers, they recover together.
  • The channels[] snapshots in session_started and resume_ok follow the same rules.

Known limitations:

  • When the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
  • Speakers are determined by channel labeling; merging and reassigning speakers through the speaker APIs do not apply (same as the common multi-channel rule), so this cannot be corrected afterward.
  • Retroactive transcription after resuming from a pause covers the whole last 60 seconds, for all channels together, with speakers labeled (see pause).

Multi-channel start-specific errors:

Error CodeHTTP StatusDescriptionRecommended Action
invalid_recognition_mode400Multi-channel is not enabled for this environmentContact the platform to enable multi-channel
channel_mode_required400channel_mode is missingProvide channel_mode (per_channel or shared)
invalid_channel_mode400channel_mode is not per_channel or shared; also returned when shared is not enabled in this environment (explained in details.message)Use per_channel or shared; if not enabled, use per_channel
channel_language_not_allowed400A channel carries transcription_languages in shared mode (details carries channel_id)Remove transcription_languages from each channel and use the session-level setting
channels_required400channels is missingProvide 1–8 channel configurations (including the main speaker)
too_many_channels400The channel count exceeds the limit (details carries max and received)Reduce the number of channels
invalid_channel_id400channel_id is out of range (1–8) or duplicatedUse unique numbers within 1–8
channel_language_required400A channel does not specify exactly one transcription language (details carries channel_id)Provide exactly 1 language in each channel's transcription_languages
channel_language_mismatch400The union of channel languages does not match transcription_languages (details lists both sets)Reconcile the two language lists
multichannel_requires_pcm400audio_format is not pcmUse pcm (16kHz / 16-bit / mono)
multichannel_tts_not_allowed400Multi-channel does not support speech synthesis in this first releaseDisable tts_enabled
multichannel_broadcast_not_allowed400Broadcast does not support multi-channel modeUse another recognition mode for broadcast
invalid_parameter400Combined with speaker_diarization, type=conversation, or a malformed speaker_name (details.field indicates the field)Fix the parameters according to details

Broadcast Mode Description (type: "broadcast")

In broadcast mode, the language settings are automatically obtained from the broadcast channel settings and do not need to be sent in the WebSocket message.

Required parameters:

ParameterTypeDescription
typestringMust be "broadcast"
broadcast_tokenstringBroadcast token (obtained after creating the broadcast via the REST API)
audio_formatstringAudio format (pcm or webm)

Optional parameters (override the broadcast channel settings):

ParameterTypeDescription
tts_configobjectMulti-language TTS settings (overrides the settings from creation time)
summary_templatestringSummary template slug (overrides the settings from creation time; if not provided, the broadcast channel default is used)

Auto-configured parameters (can be omitted):

  • transcription_languages: read automatically from the broadcast settings
  • translation_languages: read automatically from the broadcast settings
  • realtime_translation: always enabled in broadcast mode, and false is treated as true; translation is always billed at the real-time translation rate
  • summary_template: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)
  • summary_language: read automatically from the broadcast settings (the value passed via WebSocket takes precedence)

Both the host and viewers receive interim translations: a sentence's translation first arrives with is_final: false, followed by the finalized version with is_final: true. Overwrite the display by sid plus language; to show only finalized translations, skip is_final: false.

The recording name does not reuse the channel name. If start omits name, a name such as Broadcast #1 is generated; naming otherwise works the same as for other types, see Recording Name Rules.

Broadcast Phase Descriptions:

broadcast_phaseDescriptionBehavior
live (default)Live phaseSTT/translation results are broadcast to viewers and written to the transcript.
standbyStandby phaseSTT/translation results go only to the host; viewers see the standby_message.

Standby phase purpose: Lets the host warm up STT/translation before going live, confirm that equipment is working, and then switch to the live phase.

The standby phase has a time limit (30 minutes by default): when the accumulated standby time reaches the limit, the session ends automatically, with a warning about 2 minutes before. See Standby Time Limit.

broadcast_phase accepts only lowercase standby and live: an empty string is treated as live, and any other value is rejected (invalid_parameter).

Broadcast Mode Request Example:

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "broadcast",
    "broadcast_token": "a3f9",
    "audio_format": "pcm"
  }
}

Broadcast Mode Request Example (Standby Phase + Override Summary Template):

{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "broadcast",
    "broadcast_token": "a3f9",
    "audio_format": "pcm",
    "broadcast_phase": "standby",
    "standby_message": "The talk is about to begin, please wait...",
    "summary_template": "lecture"
  }
}

Summary template priority: The value passed in the WebSocket start > the default set when the broadcast channel was created. If neither is set, no summary is automatically generated.

Broadcast Mode TTS Settings (tts_config):

Use the tts_config parameter to specify which translation languages should produce TTS audio for viewers.

tts_config FieldTypeDescription
voicestringTTS voice name
speaking_ratenumberSpeaking rate (0.5–2.0, default 1.0). Values outside the range are adjusted to the nearest bound
{
  "type": "voice-translation",
  "data": {
    "action": "start",
    "type": "broadcast",
    "broadcast_token": "a3f9",
    "audio_format": "pcm",
    "tts_config": {
      "en-US": {
        "voice": "en-US-JennyNeural",
        "speaking_rate": 1.0
      },
      "ja-JP": {
        "voice": "ja-JP-NanamiNeural",
        "speaking_rate": 1.0
      }
    }
  }
}

Note:

  • TTS languages must be valid languages in translation_languages; invalid languages are automatically ignored.
  • The host (WebSocket) does not receive TTS audio; only SSE viewers receive the tts_ready event.
  • TTS is sent only during the live phase; nothing is sent during the standby phase.

TTS Playback Mode Descriptions

ModeDescriptionBehavior
syncSynchronous mode (default)Automatically plays the latest is_final=true translated sentence; if the previous sentence is still playing, it enters the queue and waits.
asyncAsynchronous mode (manual control)The user can choose any translated sentence for TTS, controlled with the tts_play command.

Success Response

After a successful start, a session_started event is returned containing complete session initialization info. For real-time recording, the first minute is deducted first, and the event is returned once that deduction completes (see Pricing).

General recordings (transcribe / conversation / record):

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}

Broadcast mode (broadcast):

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "broadcast",
    "recognition_mode": "multi_speaker",
    "phase": "standby",
    "viewer_count": 0,
    "queue_count": 0,
    "peak_viewers": 0,
    "total_viewers": 0,
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}
FieldTypeDescription
session_idstringSession ID
task_idstringTask ID (can be used for subsequent API queries)
recording_typestringRecording type: transcribe, conversation, record, broadcast
recognition_modestringRecognition mode: single, multi_speaker
phasestringBroadcast phase: standby or live (broadcast mode only)
viewer_countintCurrent number of online viewers (broadcast mode only)
queue_countintNumber of viewers waiting in the queue (broadcast mode only)
peak_viewersintPeak number of viewers for this broadcast (broadcast mode only)
total_viewersintTotal cumulative number of viewers who have connected (broadcast mode only)
messagestringStatus description message

Multi-channel mode: the top level of data additionally carries channel_mode and channels[]; see "Multi-Channel Mode Description" above for an example and field descriptions.

Automatic End After a Long Silence

A recording that goes a certain length of time without recognizing any text (including interim results that are not final yet) ends automatically, so a recording someone forgot to stop does not keep being billed.

  • The default threshold is 900 seconds (15 minutes) and can be changed per session with silenceTimeoutSeconds. About 2 minutes before the end, a warning is sent once.
  • Any recognized text restarts the count from 0; resume and start_speaking also restart it.
  • The count does not run in these cases:
    • While paused: the count restarts from 0 after resuming. A paused recording is still billed; to stop billing, end the recording.
    • While a disconnected session waits to be resumed: after resuming, the count continues from where it was.
    • Broadcasts (type: "broadcast"): never end for lack of speech. The standby phase has its own time limit; see the Broadcast Guide.
  • Multi-channel: counts only when no channel recognizes any text; a single silent channel does not end the recording.
  • Conversation manual mode: speech recognized while the speak button is not pressed also restarts the count.

Values of silenceTimeoutSeconds (top level of the start data, optional)

ValueEffect
Omitted, or nullUses the default threshold
0This session never ends automatically for lack of speech; suited to sessions kept open for long periods during which nobody may speak
Integer from 60 to 86400The threshold in seconds for this session
Any other valuestart is rejected with the error code invalid_parameter (details.field is silenceTimeoutSeconds) and the recording does not start
  • The value must be a JSON integer: strings (for example "900"), decimal notations (for example 900.0 or 1e3), and booleans are all rejected.
  • The parameter name is silenceTimeoutSeconds; sending silence_timeout_seconds is rejected (details.field is silence_timeout_seconds), not ignored.
  • A valid value has no effect on broadcasts, but an invalid value is still rejected.
  • With a threshold of 120 seconds or less, no warning is sent; the recording simply ends when the time is up.
  • Resuming a disconnected session keeps the original setting; send it again when starting a new recording.

Events you receive

  1. stt_silence_warning (an error event with severity: "warning"): a warning; the recording continues. details.silenceSeconds is how long the silence has lasted and details.remainingSeconds is how many seconds remain. Do not treat it as the end of the recording; recognized text or resuming the recording restarts the count.
  2. stt_silence_timeout (severity: "fatal"): the recording has ended automatically. details.silence_seconds is the threshold in seconds.
  3. Then status: "ended" and task_complete, in that order, just as when the client sends stop: the recording is saved and summarized as usual, and billing runs until the end.
{
  "type": "error",
  "data": {
    "error_code": "stt_silence_warning",
    "severity": "warning",
    "message": "No speech detected for a while; the recording will end automatically soon",
    "context": "stt",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-09-25T10:28:00.000Z",
    "details": {
      "silenceSeconds": 780,
      "remainingSeconds": 120
    }
  }
}

After stt_silence_timeout, stop sending audio. Each audio message that arrives afterwards gets a session_not_started reply, which you can ignore.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
missing_transcription_languages400No language parameter providedMake sure the request includes transcription_languages
invalid_transcription_language400Invalid language codeConfirm the language code format is correct (such as zh-TW)
too_many_languages400Number of languages exceeds the limit (details.max may be the system limit or the plan's cap on simultaneously recognized transcription languages)Up to 10 transcription languages and 12 translation languages; on an unlimited plan, adjust per details.max
invalid_recording_type400Invalid recording typeUse a valid type value
audio_format_unsupported400audio_format is not supported; details.supported_formats lists the accepted valuesSwitch to a supported audio format
invalid_summary_template400Invalid summary templateConfirm the template identifier is correct
stt_init_failed503Service initialization failedRetry later
auth_insufficient_credit402Insufficient creditTop up your credit balance
auth_quota_exceeded402Available credits are insufficient; the recording did not start (less than one minute for real-time recording; the connection is not closed; details.remaining_budget is the available credit as of the most recent settlement, and details.budget_scope states whose credit it is)Top up and start again
daily_limit_reached—Usage has reached the plan's limit; the recording did not startAvailable again after the plan's reset (daily limits reset the next day)
auth_service_error500Service temporarily unavailable; the recording did not start (the connection is not closed)start again later
plan_feature_not_allowed403The unlimited plan does not include a feature enabled in the request (the connection is not closed)Use features included in your plan or upgrade the plan; query GET /api/v1/me/plan for the plan contents
concurrency_limit_reached—This API Key has reached its concurrent recording limit (the connection is not closed)start again after another recording ends
service_shutdown—The service is shutting down; the recording did not start, and the connection is then closedReconnect later and start again
tts_init_failed503TTS service initialization failedRetry later
tts_invalid_language400TTS language not in the translation languagesConfirm tts_language is in translation_languages
broadcast_token_required400Broadcast mode requires a tokenThe broadcast type must provide a broadcast_token
broadcast_token_invalid401Invalid broadcast tokenConfirm the token is correct and not expired
broadcast_not_ready503Broadcast service not yet startedRetry later
summary_invalid_mode400summary_mode is not builtin / customUse a valid mode
summary_mode_field_mismatch400The mode and field combination does not match (a required field is missing / a forbidden field was included)Adjust fields per the mode rules
summary_prompt_too_long400summary_prompt exceeds 3000 charactersShorten the custom prompt
summary_prompt_slug_too_long400summary_prompt_slug exceeds 64 charactersShorten the identifier
summary_prompt_slug_invalid400summary_prompt_slug contains control characters (\n / \r / \t / \0, etc.)Remove the control characters
invalid_parameter400Invalid parameter, for example an invalid value or name for silenceTimeoutSeconds, an invalid broadcast_phase value, broadcast_token sent with a type other than broadcast, name longer than 60 characters, summary_language longer than 20 characters, or a value outside the accepted list for options.speaking_speed, options.profanity_handling, conversation_mode, or tts_mode (details.valid_values lists the accepted values)Fix the parameter indicated by details.field

Voice Translation - config (Set Terminology / Correction Rules)

Description

Before or during recording, pass in terminology, fuzzy-word correction rules, and translation dictionary settings. These settings improve STT accuracy, fix homophone errors, and ensure translation consistency.

Terminology also drives homophone correction: When terminology is provided, the terms become the reference for homophone matching — any span in the transcript that sounds the same but is written differently is corrected back to the spelling of the term. Providing terminology alone is therefore enough to get correction; you do not need to list possible misspellings by hand.

Limits

Per-block entry limits, length limits, and the matching error codes are collected in Terminology Guide → Limits.

The numbers are defaults; the limit actually in force can be tuned per environment — always treat max in the error response's details as authoritative rather than hard-coding the numbers.

Use Cases

  • Pass in professional terminology (Phrase List) before recording starts
  • Set fuzzy-word correction rules (homophone correction) - optional; terminology itself already provides homophone correction
  • Set a translation dictionary (ensure consistent terminology translation)

Timing

Setting TypeRecommended TimingUpdate During Recording
TerminologyBefore or during startSupported (from the next utterance)
Fuzzy-word correctionBefore or during startSupported (from the next utterance)
Translation dictionaryBefore or during startSupported (from the next utterance)

Note: When you update terminology, fuzzy-word correction, or the translation dictionary during recording, the new settings take effect from the next utterance after they are sent; the utterance currently being recognized is not guaranteed to be covered. No reconnection is needed. When terminology is updated, the response includes a terminology_effective: "next_turn" field as a hint.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value config
terminologyobjectNoTerminology settings
fuzzy_correctionobjectNoFuzzy-word correction rules
translation_dictobjectNoTranslation dictionary

Note: At least one setting item must be provided.

Note: All or nothing: all three blocks are validated before any of them is applied. If any block fails validation, an error is returned and none of the three blocks is applied — the settings stay as they were.

For example, sending a valid terminology list together with more than 3000 translation-dictionary entries returns config_too_many_dict_entries, and the terminology does not take effect either. Fix the problem and resend the complete config; all three blocks replace their previous value wholesale, so resending does not stack on top of earlier settings.

Terminology Format (terminology)

Keyed by language code, with an array of terms as the value:

{
  "zh-TW": [
    { "term": "語者分離" },
    { "term": "WebSocket" }
  ],
  "en-US": [
    { "term": "diarization" }
  ]
}
FieldTypeRequiredDescription
termstringYesThe term (max 100 characters)

Limit: Up to 500 terms across all languages combined in one config — not 500 per language. Exceeding the limit returns config_too_many_entries (with count and max in details).

Language scope: Terms apply only to recognition in the language they are registered under. Single-language situations — each multi-channel track, speaker diarization, and file import — use only that language's terms; multi-language transcription and conversation mode use the terms for the languages declared for the session. Terms registered under a language not used in this session have no effect, but still count toward the 500 combined total above.

These numbers are defaults: the limit actually in force can be tuned per environment; always treat max in the error response's details as authoritative.

Fuzzy-Word Correction Format (fuzzy_correction)

Note: This field usually does not need to be set manually — terminology already corrects misspellings that sound the same or nearly the same. Use it in these three cases:

  1. The misspelling is itself an ordinary word and is therefore blocked by common-word protection — 晶圓 heard as 金元, for example
  2. The misspelling sounds very different from the correct term, such as a foreign brand name recognized as a phonetically unrelated word
  3. Misspellings in Japanese, Korean or English, which do not participate in homophone matching

When correct is Chinese, it also becomes a reference for homophone matching — homophone misspellings not listed in incorrect are corrected to correct as well. case_insensitive applies only to the literal matching of incorrect; it does not affect homophone matching.

Keyed by language code, with an array of correction rules as the value:

{
  "zh-TW": [
    { "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] },
    { "correct": "IPEVO", "incorrect": ["ltfo"], "case_insensitive": true }
  ]
}
FieldTypeRequiredDescription
correctstringYesThe correct word
incorrectstring[]ConditionalList of incorrect variants, each up to 200 characters. Can be omitted for Chinese terms (see below); required otherwise — an empty array returns config_invalid_entry (reason: "empty")
case_insensitivebooleanNoWhether this rule's variants match regardless of case (defaults to false = exact-case matching)

Supplying only the correct term: when correct is Chinese (contains Han characters), incorrect may be omitted entirely — the system matches by pronunciation, and spellings in the transcript that sound the same or nearly the same are corrected back to correct.

{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }

No misspellings need to be listed above: 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected. Only spellings that sound quite different (愛自動, say) or have a different number of syllables (愛松) still need to be listed in incorrect.

Note: Both conditions must hold: the language must be Chinese (zh-TW, zh-CN, zh-HK and so on) and correct must contain Han characters. Otherwise incorrect remains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took. List misspellings explicitly for Japanese, Korean and English.

Case sensitivity: case_insensitive is optional and defaults to false (exact-case matching). When set to true, every incorrect variant in that rule matches regardless of case. The flag is per rule — the same correct term can be split across several rules with different settings, for example making variants that cannot collide with ordinary words case-insensitive while keeping variants that could hit a personal name exact. It has no effect on Chinese rules (Chinese has no letter case).

Note: Enabling it widens the false-positive surface: if ivo is case-insensitive, the personal name Ivo is replaced too.

When the same incorrect variant appears in more than one rule: collisions are resolved on incorrect (the variant), not on correct.

  • When several rules point at different correct terms, which one actually takes effect is not guaranteed; do not rely on any ordering (registration order included)
  • The case flag resolves to strict wins — if any rule leaves case_insensitive off, that variant is matched with exact case

"Strict wins" is a deliberately conservative choice: it prevents a permissive rule elsewhere in your vocabulary from silently loosening a brand-name rule you explicitly set to strict.

Splitting one correct term across several rules is therefore safe, as long as their incorrect variants do not overlap. But if your data contains the same incorrect variant mapped to different correct terms, the later one is dropped with no warning — check for duplicate variants before sending.

How Terminology Participates in Homophone Correction

Once terminology is provided, the terms become the reference for homophone matching. Any span in the transcript that sounds the same or nearly the same but is written differently is corrected back to the spelling of the term:

TermAppears in transcript asCorrected to
紡拓會訪拓會紡拓會
語者分離語這分離, 與者分離語者分離
晶圓晶園晶圓

You do not need to list the possible misspellings in advance — matching is based on pronunciation, so coverage is not limited by how many variants you can think of.

Mixed Chinese-English terms: For a term like CVD製程, matching applies only to the Chinese portion; the Latin portion is left unchanged.

Applicable languages: Homophone matching applies to Chinese only (Traditional and Simplified are interchangeable, since they share pronunciation). Terms in Japanese, Korean or English do not participate in homophone matching — list their misspellings explicitly with fuzzy_correction.

Matching range: Both identical and near-identical pronunciations are covered, including accent differences such as jin vs jing (final -n vs -ng) and retroflex vs non-retroflex initials. For the term 晶圓廠, for example, 金圓廠 in the transcript is corrected.

Common-word protection: If the span in the transcript is itself an ordinary word (金元 or 反案, say), it is not changed even when it shares a pronunciation with a term — this prevents normal sentences from being altered. To force such a correction, list the misspelling explicitly with fuzzy_correction: explicitly listed misspellings are not subject to common-word protection.

Limit: Up to 4000 rules across all languages combined in one config. Exceeding it returns config_too_many_entries (with field: "fuzzy_correction", count and max in details).

This number is a default: the limit actually in force can be tuned per environment, so always treat max in the error response's details as authoritative. Do not hard-code the numbers in your integration — if you want to check before sending, read max and fill it back in. When you do hit a limit, details carries both count (what you sent) and max (the limit in force).

Language scope: A correction rule applies only to sentences in the language it is registered under — a rule under zh-TW will not alter an English sentence. When the sentence language cannot be determined, all rules are applied as a fallback (better to over-apply than to skip the sentence entirely). Homophone matching driven by terminology likewise follows the language each term is registered under.

Translation Dictionary Format (translation_dict)

Group entries by language code, giving each language its own dictionary:

{
  "en-US": [
    { "source": "語者分離", "target": "Speaker Diarization" },
    { "source": "晶圓", "target": "wafer", "case_sensitive": true }
  ],
  "ja-JP": [
    { "source": "語者分離", "target": "話者分離" }
  ]
}
FieldTypeRequiredDescription
(top-level key)stringYesTarget language code
sourcestringYesThe source word (in the STT language), up to 200 characters
targetstringYesThe required translation for this language, up to 200 characters
case_sensitivebooleanNoWhether the entry applies only on an exact-case match (defaults to false = case-insensitive)

Limit: Up to 3000 entries per language. Exceeding the limit returns config_too_many_dict_entries; the details in the response identify which language exceeded it.

Note: The entry count directly affects translation workload and cost. A single translation only carries entries whose source word actually appears in that piece of text, up to 100 of them; beyond that, longer source words are kept first. Also, the more entries there are, the smaller the share that is reliably honored — an inherent limit that a higher cap does not change.

The previous format is still supported: the earlier array-of-entries format ([{ "source": ..., "translations": { "language code": ... } }]) is still accepted, with identical content and behavior. Existing integrations need no changes. The same dictionary sent in either format produces identical results.

On resume, the translation_dict returned in resume_ok matches the format you last sent — send the old format and you get the old format back; send the new one and you get the new one.

Note: Clearing the whole dictionary is not supported: an empty object is treated as not sending this setting at all. Clearing a single language is supported — send that language an empty array (e.g. {"en-US": []}).

Case sensitivity: case_sensitive is optional and defaults to false (case-insensitive). When set to true, the entry applies only where the source text matches source exactly, including case. The flag is per entry.

Note: The translation dictionary guides the model through prompting rather than literal substitution, so it is best-effort, not deterministic — the case flag is likewise a hint and is not guaranteed to be honored. Use fuzzy_correction when you need deterministic replacement.

Case-Flag Comparison

fuzzy_correction and translation_dict each have a case switch. Their field names are opposites, and so is the behavior their default value produces:

BlockFieldDefaultDefault behavior
fuzzy_correctioncase_insensitivefalseStrict (case-sensitive)
translation_dictcase_sensitivefalsePermissive (case-insensitive)

Both default to false, yet one means strict and the other means permissive. Do not share a single variable between them, and do not mirror one value onto the other — getting it wrong produces no error at all, only matching behavior opposite to what you intended.

{
  "type": "voice-translation",
  "data": {
    "action": "config",
    "terminology": {
      "zh-TW": [
        { "term": "語者分離" },
        { "term": "CVD製程" },
        { "term": "wafer良率" }
      ]
    }
  }
}

Request Example (Full Settings, Including Manual Correction Rules)

{
  "type": "voice-translation",
  "data": {
    "action": "config",
    "terminology": {
      "zh-TW": [
        { "term": "語者分離" },
        { "term": "即時轉錄" }
      ]
    },
    "fuzzy_correction": {
      "zh-TW": [
        { "correct": "語者分離", "incorrect": ["語這分離", "語者分力"] }
      ]
    },
    "translation_dict": {
      "en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }]
    }
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "config_updated",
    "updated": ["terminology", "fuzzy_correction", "translation_dict"],
    "message": "Settings updated"
  }
}
FieldTypeDescription
updatedstring[]The setting types that were updated
messagestringStatus message

Error Responses

Important: Your client must also listen for type: "error" messages — do not wait only for config_updated.

When the server rejects a configuration it sends a type: "error" message and notconfig_updated. An integration that waits only for config_updated will hang until its own timeout and appear as "the server never responded", even though the error was delivered and data.error_code already states the reason.

Error CodeHTTP StatusDescriptionRecommended Action
config_empty400No configuration provided. Note: An empty object {} does not count as "provided" — sending {"terminology": {}, "fuzzy_correction": {}, "translation_dict": {}} triggers this errorProvide at least one setting that actually has content. To clear a language's glossary, send {"lang": []} (for example {"zh-TW": []})
config_term_too_long400Term exceeds 100 charactersShorten the term length
config_too_many_entries400More than 500 terms, or more than 4000 fuzzy correction rules (both across all languages combined)Remove terms or correction rules
config_too_many_dict_entries400Translation dictionary exceeds 3000 entries for a single language (details.language identifies which one)Reduce the dictionary entries for that language
config_invalid_entry400A glossary entry has an invalid field (details carries language, index, field, and reason for locating it; depending on the case it may also carry variant_index, max_length, or count/max)Fix the entry at the location given in details

Voice Translation - audio (Send Audio)

Description

Sends audio data to the server for speech recognition. The audio must be Base64-encoded before sending.

Use Cases

  • Continuously sending microphone audio
  • Sending recorded audio segments

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value audio
payloadstringYesBase64-encoded audio data
channel_idintConditionalIn multi-channel mode (recognition_mode=multi_channel), required on every frame: the source channel number of this audio frame. Omitting it returns channel_id_required; an unknown or removed number returns unknown_channel_id. Outside multi-channel mode this field is always ignored (backward compatible)

Audio Format Requirements

PCM format (default):

ItemSpecification
FormatPCM (raw audio)
Sample rate16000 Hz
Bit depth16-bit
ChannelsMono
Byte orderLittle-endian
Transport encodingBase64

WebM/Opus format:

ItemSpecification
FormatWebM container + Opus codec
Sample rateAny (the server converts automatically)
ChannelsMono or Stereo (the server converts automatically)
Transport encodingBase64

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "audio",
    "payload": "Base64-encoded PCM audio data"
  }
}

Request Example (Multi-Channel Mode)

In multi-channel mode every frame must include channel_id:

{
  "type": "voice-translation",
  "data": {
    "action": "audio",
    "channel_id": 2,
    "payload": "Base64-encoded PCM audio data"
  }
}

Multi-channel mode notes: we recommend sending one frame every 100ms; every channel should keep sending audio even while silent. Multi-channel supports only the pcm format (16kHz / 16-bit / mono / little-endian).

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400Speech recognition has not startedCall the start action first
audio_invalid_format400Invalid audio data formatConfirm the Base64 encoding is correct
audio_decode_failed400Audio decoding failedConfirm the audio format is correct. The recording continues; for WebM, send a new container (with its header) to recover. Time that cannot be decoded is not billed
channel_id_required400An audio frame in multi-channel mode is missing channel_idInclude the source channel number on every frame
unknown_channel_id400Unknown or removed channel_idUse a valid number declared in start or added via add_channel

Voice Translation - pause (Pause Translation)

Description

Pauses speech recognition processing. Audio received during the pause is buffered and continues to be processed after resuming. Two-way translation is an exception: audio received while paused is not kept; for a sentence in progress at the moment of pausing, the server first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds), sends it with is_final: true, and only then returns status: "paused". Billing continues while paused; see Pricing.

Use Cases

  • The user steps away temporarily
  • You need to pause recording

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "pause"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "paused",
    "message": "Speech recognition paused"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400Speech recognition has not startedCall start first
session_already_paused400Already pausedYou can ignore this error

Multi-channel mode: while paused, every channel's audio file is still saved as usual, but no text is produced (no transcript is generated). Adding or disabling channels, changing a channel's language, or adjusting the speaking speed while paused returns channel_action_while_paused — call resume first, then operate.


Voice Translation - resume (Resume Translation)

Description

Resumes paused speech recognition processing.

Use Cases

  • The user returns to continue
  • You need to continue recording

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "resume"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "live",
    "message": "Speech recognition resumed"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400Speech recognition has not startedCall start first
session_not_paused400Not pausedYou can ignore this error

Multi-channel mode: after resuming, speech from the paused period is transcribed retroactively (timestamps land at the moment the words were actually spoken, not after the resume). Retroactive transcription has a limit: a trailing 60 seconds shared across the whole session (under per_channel, split evenly across channels when there are several; under shared, the whole last 60 seconds, for all channels together, with speakers labeled); anything earlier is kept only in the audio file and does not enter the transcript. A sentence cut off at the moment of pausing may not appear in the transcript (same as single-stream behavior); when it is indeed not kept, resuming also sends segment_discarded (reason: "resumed") to identify it. On resume, each channel prepares recognition again and sends its own channel_status event (reason: "resumed").


Voice Translation - stop (Stop Translation)

Description

Stops speech recognition and ends the session. The server first waits for the recognizer to finish the last sentence (usually about 1 second, at most about 3 seconds) so that it is included in the transcript; if it does not arrive in time, the last interim result shown is used as that sentence. The system automatically uploads the audio file and transcript, and generates a summary (if configured; when the available credits cannot cover the summary fee, no summary is generated and summary_error is sent instead).

Use Cases

  • The meeting ends
  • Recording is complete

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "stop"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "ended",
    "message": "Speech recognition stopped"
  }
}

Task Complete Event

This event is sent after the audio file and transcript have been uploaded:

{
  "type": "voice-translation",
  "data": {
    "action": "task_complete",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "message": "Task processing complete"
  }
}
FieldTypeDescription
task_idstringRecording UUID, can be used for subsequent API queries
noAudiobooleanPresent only when no audio was received during the whole recording, and always true; see task_complete

Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
session_not_started400The session has not started, or this recording has already ended (for example, stop sent twice)Call start first if it has not started; if this was a duplicate stop, the error can be ignored

Sending stop a second time returns this error rather than another success response. task_complete is sent only once, after the first successful stop, when the audio and transcript have finished uploading.


Voice Translation - retranslate (Retranslate)

Description

Retranslates a specified sentence, useful when the original text has been corrected and the translation needs to be updated.

Use Cases

  • The user edits the original text and the translation needs updating
  • Correcting recognition errors

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value retranslate
sidintYesThe sentence number to retranslate
translation_languagesstring[]YesArray of translation language codes
textstringYesThe original text to translate (the user-corrected text)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "retranslate",
    "sid": 1,
    "translation_languages": ["en-US"],
    "text": "The user-corrected original text"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "The new translation result",
        "is_final": true,
        "is_retranslation": true
      }
    }
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_data422No sid providedInclude sid
record_translation_not_allowed400Recording-only sessions do not support translationUse the transcribe type
retranslate_session_not_active400The session is not started or has endedConfirm the session state
retranslate_no_target_lang400No target language providedProvide translation_languages
retranslate_no_text400No text to translate providedProvide the text parameter
retranslate_llm_not_ready503The translation service is not readyRetry later
retranslate_llm_failed500Translation service failedRetry later

If the translation service returns a specific failure code (for example llm_content_filtered when the content cannot be translated), that code is returned as-is instead of being wrapped in retranslate_llm_failed.


Voice Translation - switch_language (Switch Language)

Description

Switches or adjusts translation languages while real-time translation is in progress. The behavior varies by recording type and the number of translation languages:

  • General mode, single language (translation_languages has 1 entry): replaces the translation target language and automatically batch-retranslates all already-translated sentences.
  • General mode, multiple languages (2 or more entries, v1.6.7): redefined as "add or remove a single language". The op parameter (add / remove) is required; omitting it returns a switch_language_op_required error. add automatically backfills existing sentences (same response sequence as single-language replace); remove returns a translation_language_removed event and keeps existing translations. See the WebSocket reference
  • Two-way mode (conversation): switches the STT source language (spoken language); the translation target automatically switches to the other language.
  • Multi-channel mode (multi_channel, v1.10.0): not supported — always returns multichannel_switch_language_not_allowed (including op: add / op: remove). Languages are bound to channels; use set_channel_language instead.

Use Cases

  • Switching the translation target language
  • A change in language needs mid-meeting
  • Adding / removing a translation language mid-session in multi-language sessions (v1.6.7)

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value switch_language
translation_languagesstring[]ConditionalArray of translation language codes (required in general mode; only the first element is used as the operation target)
opstringConditionalv1.6.7 multi-language operation: add or remove. Required for multi-language sessions
transcription_languagesstring[]ConditionalThe target language to switch to (two-way mode; if omitted, automatically toggles to the other language)

Request Example (General Mode, Single-Language Replace)

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "translation_languages": ["ja-JP"]
  }
}

Request Example (Multi-Language Add / Remove, v1.6.7)

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "op": "add",
    "translation_languages": ["de-DE"]
  }
}
{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "op": "remove",
    "translation_languages": ["ko-KR"]
  }
}

Request Example (Two-Way Mode)

Specify the target to switch to:

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language",
    "transcription_languages": ["en-US"]
  }
}

Automatic toggle (no parameters):

{
  "type": "voice-translation",
  "data": {
    "action": "switch_language"
  }
}

Two-Way Mode Special Behavior:

  • Two-way mode uses automatic language detection and usually does not require manually switching the language.
  • switch_language only updates the internal preference state.
  • After a successful switch, a language_switched event is returned (not a language_switch_start/done sequence).
  • Switching to the same language returns a conversation_same_language warning.

Response Sequence (General Mode)

After switching the language, you receive the following events in order:

  1. language_switch_start: notifies that the switch has begun
{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_start",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "total_segments": 15
  }
}
  1. batch_retranslation (multiple): returns retranslation results sentence by sentence
{
  "type": "voice-translation",
  "data": {
    "action": "batch_retranslation",
    "sid": 3,
    "translations": {
      "ja-JP": {
        "sid": 3,
        "text": "今日はプロジェクトの進捗について話し合いましょう",
        "is_final": true,
        "is_retranslation": true
      }
    }
  }
}
  1. language_switch_done: notifies that the switch is complete
{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_done",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "success_count": 15,
    "failed_count": 0
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
switch_language_no_target400No target language providedProvide translation_languages
switch_language_in_progress400The previous switch is not yet completeWait for the switch to complete
switch_language_same_target400The target language is the same as the current oneYou can ignore this error
conversation_requires_two_languages400Two-way mode requires exactly two languagesConfirm transcription_languages has 2
conversation_languages_identical400The two two-way languages cannot be the sameProvide two different languages
conversation_invalid_language400Invalid two-way languageConfirm the language is in transcription_languages
conversation_same_language400Already the current languageYou can ignore this warning
multichannel_switch_language_not_allowed400Multi-channel mode does not support switch_languageUse set_channel_language to change a single channel's language

Voice Translation - set_name (Set Recording Name)

Description

Sets the name while recording is in progress. After it is set, this name is used when the recording ends and will not be auto-generated.

Tip: You can also set an initial default name via the name parameter at start, but that name may still be overridden by the system when the session ends. If you need a fixed name, use set_name.

Use Cases

  • Customizing the recording title after recording starts
  • Overriding an auto-generated name or a previously set name

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_name
namestringYesRecording name (max 60 chars)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_name",
    "name": "Product Planning Meeting"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "event": "name_set",
    "name": "Product Planning Meeting",
    "message": "Recording name updated"
  }
}

Compatibility note: For backward compatibility, the set_name success response keeps action: "status" (unchanged). New clients should identify a successful set_name via event: "name_set" (together with the name field). Relying on action: "status" to detect set_name success is deprecated and may be removed in a future version.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
set_name_empty400Recording name is emptyProvide a non-empty name
set_name_too_long400Recording name exceeds the length limit (>60 chars); the response details includes max_lengthShorten the name (≤60 characters)

Voice Translation - rename_speaker (Globally Rename a Speaker)

Description

In multi-speaker diarization mode (multi_speaker), globally renames a speaker. All sentences using that speaker ID are updated in sync.

Also available in multi-channel mode (multi_channel): speaker IDs use the format channel_{N} (such as channel_1); speakers are registered at start / add_channel, so a channel's speaker can be renamed before they say anything. After renaming, the speaker_label in result events and the transcript both reflect the new label.

Use Cases

  • Changing a system-assigned speaker ID (such as Guest-1) to a meaningful name (such as Manager Wang)
  • Naming a newly recognized speaker during a meeting

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value rename_speaker
speaker_idstringYesThe original speaker ID (such as Guest-1); the current display label is also accepted for consecutive renaming; max 100 characters
new_labelstringYesThe new display label; max 100 characters, must not contain control characters (\x00-\x1F, \x7F) or line breaks

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "rename_speaker",
    "speaker_id": "Guest-1",
    "new_label": "Manager Wang"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_renamed",
    "speaker_id": "Guest-1",
    "new_label": "Manager Wang",
    "affected_sids": [1, 3, 5, 8]
  }
}
FieldTypeDescription
speaker_idstringThe resolved original speaker ID (even if the input was a display label, the event returns the original ID)
new_labelstringThe new display label
affected_sidsint[]The list of affected sentence numbers

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
speaker_not_found422The specified speaker was not foundConfirm the speaker_id or display label exists
speaker_name_empty422new_label is emptyProvide a valid label
speaker_name_duplicate422The display label is already in useUse a different label, or first change the conflicting speaker
session_not_started400Speech recognition has not startedCall start first

Voice Translation - reassign_speaker (Change the Speaker of a Single Sentence)

Description

Changes the speaker identity of a specific sentence, assigning the sentence to an existing speaker.

Not applicable in multi-channel mode (multi_channel): speaker identity is determined by the physical channel (each sentence's speaker_id is its source channel), so reassigning a sentence to another channel is not supported.

Use Cases

  • Correcting a speaker identity that the system recognized incorrectly
  • Reassigning a sentence to another known speaker

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value reassign_speaker
sidintYesThe sentence number to change
target_speaker_idstringYesThe target speaker's original ID (taken from init_sentence.speaker_id; reassign does not accept display labels)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "reassign_speaker",
    "sid": 5,
    "target_speaker_id": "Guest-2"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_reassigned",
    "sid": 5,
    "old_speaker_id": "Guest-1",
    "new_speaker_id": "Guest-2",
    "new_speaker_label": "Lee Hsiao-hua"
  }
}
FieldTypeDescription
sidintThe changed sentence number
old_speaker_idstringThe original speaker ID
new_speaker_idstringThe new original speaker ID
new_speaker_labelstringThe new speaker display label (after applying speaker_aliases; equals new_speaker_id when no alias exists)

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
speaker_sid_not_found422The specified sentence was not foundConfirm the SID exists
speaker_not_found422The target speaker does not existUse an existing speaker ID
speaker_name_empty422The target speaker ID cannot be emptyProvide a valid speaker ID
session_not_started400Speech recognition has not startedCall start first
invalid_parameter400Creating a new speaker is not supportedUse an existing speaker ID

Voice Translation - merge_speakers (Merge Speakers)

Description

Merges all sentences of one speaker into another speaker. After the merge, future recognition results for that speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.

Not applicable in multi-channel mode (multi_channel): speaker identity is determined by the physical channel, so merging speakers across channels is not supported.

Use Cases

  • The speech recognition engine sometimes misidentifies the same person's voice as multiple speakers (for example, Guest-1 and Guest-2 are actually the same person)
  • Use this feature to merge all of Guest-2's sentences into Guest-1
  • After the merge, future Guest-2 recognition results are automatically displayed as Guest-1
  • This interception applies only to the current recognition pass: if the connection drops and recovers mid-recording, speakers are identified again and receive new IDs, and the merge does not carry over — merge again once you confirm it is the same person

Difference from reassign_speaker

FeatureScopeFuture Impact
reassign_speakerA single sentence (1 SID)None
merge_speakersAll sentences of that speakerFuture appearances of the source are also automatically converted to the target

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value merge_speakers
source_speaker_idstringYesThe speaker ID to be merged (such as Guest-2)
target_speaker_idstringYesThe merge target speaker ID (such as Guest-1)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "merge_speakers",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "speakers_merged",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1",
    "affected_sids": [3, 5, 7]
  }
}
FieldTypeDescription
source_speaker_idstringThe original ID of the merged speaker
target_speaker_idstringThe original ID of the merge target
affected_sidsnumber[]The list of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target)

To obtain the target speaker's display label, query speaker_aliases or the next init_metadata event.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
speaker_not_found422The speaker does not existConfirm the speaker ID exists
merge_speakers_same_id400The source and target speaker are the sameUse different speaker IDs
speaker_name_empty422The speaker ID cannot be emptyProvide a valid speaker ID
session_not_started400Speech recognition has not startedCall start first

Voice Translation - tts_play (Play TTS)

Description

In async mode, manually plays the TTS audio for a specified sentence.

Use Cases

  • The user selects a specific sentence for TTS playback
  • Playing multiple consecutive sentences

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value tts_play
sidintYesThe starting sentence ID
lengthintNoNumber of sentences to play (default 1, max 20)

Note: The maximum value of length is controlled by a server-side setting (default 20).

Request Example (Single Sentence)

{
  "type": "voice-translation",
  "data": {
    "action": "tts_play",
    "sid": 5
  }
}

Request Example (Multiple Sentences)

{
  "type": "voice-translation",
  "data": {
    "action": "tts_play",
    "sid": 5,
    "length": 3
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
tts_not_enabled400TTS not enabledConfirm TTS was enabled at start

A missing sentence or translation does not return error: If the starting sid does not exist, or the sentence has no translation in the target language, no error message is returned. A tts_error event is sent instead (see "Response Events"), with error set to sentence_not_found and translation_not_found respectively. When playing multiple sentences, a failing sentence is skipped and the rest still play.


Voice Translation - tts_stop (Stop TTS)

Description

Stops the TTS audio that is currently playing.

Use Cases

  • The user manually stops TTS playback
  • Stopping the current playback before switching to another sentence

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "tts_stop"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "TTS stopped"
  }
}

Voice Translation - tts_mode (Switch TTS Mode)

Description

Switches the TTS playback mode (synchronous/asynchronous) while recording is in progress.

Use Cases

  • Switching from automatic playback to manual control
  • Switching from manual control to automatic playback

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value tts_mode
tts_modestringYesMode: sync (synchronous) or async (asynchronous); only these two lowercase values are accepted

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode",
    "tts_mode": "async"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode_changed",
    "tts_mode": "async"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_data422No tts_mode provided, or a value other than sync or async. For the latter, details.field is tts_mode and details.valid_values lists the accepted values; the mode does not change and no tts_mode_changed is sentInclude sync or async

Voice Translation - set_tts (Two-Way TTS Settings)

Description

While a two-way mode (conversation) recording is in progress, dynamically toggles TTS on/off or updates the TTS voice settings. Available only in two-way mode.

Use Cases

  • Turning the TTS audio response off/on mid-conversation in two-way mode
  • Changing the TTS voice or speaking rate for a specific language

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_tts
tts_enabledbooleanNoWhether to enable two-way TTS (true / false)
tts_configobjectNoTTS settings per language; the key is the language code, and the value is {voice, speaking_rate}

Note: At least one of tts_enabled and tts_config must be provided. tts_config updates only the settings for the specified languages; unspecified languages remain unchanged.

Request Example (Disable TTS)

{
  "type": "voice-translation",
  "data": {
    "action": "set_tts",
    "tts_enabled": false
  }
}

Request Example (Update Voice Settings)

{
  "type": "voice-translation",
  "data": {
    "action": "set_tts",
    "tts_enabled": true,
    "tts_config": {
      "en-US": {
        "voice": "en-US-GuyNeural",
        "speaking_rate": 1.2
      }
    }
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "tts_updated",
    "tts_enabled": true,
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
    }
  }
}
FieldTypeDescription
tts_enabledbooleanThe current TTS enabled state
tts_configobjectThe current complete TTS settings (all languages)

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not two-way modeThis action is available only in conversation mode
session_not_started400Speech recognition has not startedCall start first

Voice Translation - start_speaking (Start Speaking / Manual Mode)

Description

In two-way manual mode (conversation_mode: "manual"), notifies the system that the user has started speaking. From this point on, audio is sent to STT for recognition, and all recognition results accumulate into a single sentence (no automatic segmentation). If it is called again while already speaking, the system first ends the previous sentence (waiting for the last part to finish, then sending its final result) and starts a new one.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value start_speaking
speakerintYesUser number (1 or 2)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "start_speaking",
    "speaker": 1
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "Speaking started"
  }
}

If it is called again while already speaking, the final result of the previous sentence is sent as usual, followed by this same status.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not two-way modeUse only under the conversation type
conversation_not_manual_mode400Not manual modeUse only in manual mode
conversation_invalid_speaker400Invalid user numberUse 1 or 2

Voice Translation - stop_speaking (Stop Speaking / Manual Mode)

Description

In two-way manual mode, notifies the system that the user has stopped speaking. The server first waits for the recognizer to finish the last part (usually about 1 second, at most about 3 seconds). The system merges the recognition results accumulated during the period into a single complete sentence and performs translation and TTS synthesis.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value stop_speaking

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "stop_speaking"
  }
}

Success Response

After stopping speaking, the system sends a complete result event (containing origin and translations):

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "origin": {
      "sid": 1,
      "language": "zh-TW",
      "text": "The complete sentence merged from all recognition during this period",
      "is_final": true,
      "speaker_id": "Speaker-1",
      "start_time": "00:05"
    },
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "The complete merged sentence from this speaking period",
        "is_final": true
      }
    }
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not two-way modeUse only under the conversation type
conversation_not_speaking400Not in a speaking stateCall start_speaking first

Voice Translation - switch_conversation_mode (Switch Conversation Mode)

Description

While two-way mode is in progress, switches between auto-detect mode (auto) and manual mode (manual). If the user is currently speaking when the switch happens, speaking is ended automatically.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value switch_conversation_mode
conversation_modestringYesThe target mode: auto or manual

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "switch_conversation_mode",
    "conversation_mode": "manual"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "conversation_mode_changed",
    "conversation_mode": "manual"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not two-way modeUse only under the conversation type
conversation_invalid_mode400Invalid conversation modeUse auto or manual

Voice Translation - set_speaker_language (Set User Language)

Description

While two-way mode is in progress, changes a specified user's language in real time. The system rebuilds the STT connection to accommodate the new language, and the translation target is also updated automatically. The transcript content before the change keeps its original language, and the timestamp continues to count without resetting.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_speaker_language
speakerintYesUser number (1 or 2)
languagestringYesThe new language code (such as ja-JP)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_speaker_language",
    "speaker": 1,
    "language": "ja-JP"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_language_changed",
    "speaker_language_map": {
      "1": "ja-JP",
      "2": "en-US"
    }
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_action400Not two-way modeUse only under the conversation type
conversation_invalid_speaker400Invalid user numberUse 1 or 2
conversation_invalid_language400No language providedInclude language
invalid_transcription_language400Invalid language codeUse a valid BCP 47 language code
session_not_started400Recording has not startedCall start first
conversation_same_language400Same as the current languageYou can ignore this warning
conversation_language_same_as_peer400The new language is the same as the other userThe two users cannot have the same language
conversation_speaking400Currently speaking, cannot change languageEnd speaking before changing
conversation_language_change_failed500Language change failed (STT rebuild failed)Retry later

Voice Translation - set_speaking_speed (Change Speaking Speed During Recording)

Description

Dynamically adjust the speaking speed (which controls the silence threshold for segmentation) while recording. The system rebuilds the STT connection to apply the new setting, causing a brief interruption in recognition (same as changing a speaker's language during recording). Supported in all recognition modes except multi-speaker (multi_speaker) — including broadcast and multi-language LID; speaking_speed is not applied in multi-speaker mode.

Multi-channel mode (multi_channel, v1.10.0): supported. The new segmentation threshold applies to all channels and takes effect channel by channel (each channel sends its own channel_status event with reason: "speaking_speed", and briefly produces no text while the new setting is applied). The whole session shares a rate limit of one adjustment per 5 seconds — adjusting too quickly returns channel_rebuild_too_frequent; while paused it returns channel_action_while_paused.

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_speaking_speed
speaking_speedstringYesvery_slow / slow / normal / fast / very_fast (see speaking_speed levels for each threshold)

Request Example

{
  "type": "voice-translation",
  "data": { "action": "set_speaking_speed", "speaking_speed": "slow" }
}

Success Response

{
  "type": "voice-translation",
  "data": { "action": "speaking_speed_changed", "speaking_speed": "slow" }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
invalid_data422Invalid speaking_speed valueUse one of the five valid values
session_not_started400Recording not startedCall start first
invalid_action400Not supported in multi-speaker modeDo not call in multi-speaker mode
set_speaking_speed_failed400STT rebuild failedRetry later
channel_rebuild_too_frequent400Multi-channel: speaking-speed adjustments are too frequent (only one per 5 seconds is accepted; details carries cooldown_seconds)Retry later
channel_action_while_paused400Multi-channel: the speaking speed cannot be adjusted while pausedCall resume first, then adjust

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "speaking_speed"). You may receive it even when set_speaking_speed_failed is returned.


Voice Translation - add_channel (Add a Channel / Multi-Channel)

Multi-channel mode only (v1.10.0)

Description

Dynamically adds a channel while a multi-channel recording (recognition_mode: "multi_channel") is in progress. The new channel begins producing text after about 4 seconds; audio sent during this period is not lost, only delayed. After a successful add, the billed channel count follows the actual number of channels from the next minute onward.

Use Cases

  • A new speaker joins mid-meeting and gets a new microphone
  • Opening channels one by one as attendance grows

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value add_channel
channelsarrayYesExactly 1 channel configuration element (same fields as channels[] in start: channel_id, speaker_name, transcription_languages; in shared mode, omit transcription_languages)

Numbers cannot be reused: once a channel_id has been used (including ones already removed via remove_channel), it can never be used again; reusing it returns channel_id_in_use. Speaker identity in the transcript is fixed at the moment each sentence is written, and reusing a number would make one speaker ID correspond to two different people.

Shared mode: the added channel shares the session's recognition and does not increase the billed channel count; its status takes on the first channel's current status directly (see Shared Mode).

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "add_channel",
    "channels": [
      { "channel_id": 4, "speaker_name": "Director Lin", "transcription_languages": ["ja-JP"] }
    ]
  }
}

Success Response

On success, a channel_status event is returned (reason: "added"). The new channel starts in preparing; once that channel produces its first text, another event with status: "ready" arrives:

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 4,
      "speaker_name": "Director Lin",
      "transcription_languages": ["ja-JP"],
      "status": "preparing"
    },
    "active_channels": 4,
    "stt_stream_count": 4,
    "reason": "added"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
not_multi_channel_session400This recording is not in multi-channel modeAvailable only with recognition_mode=multi_channel
channels_required400channels is missing or does not contain exactly 1 elementProvide exactly 1 channel configuration
invalid_channel_id400channel_id is out of range (1–8)Use a number within 1–8 that has never been used
channel_id_in_use400The number is in use or has already been removed; it cannot be reusedUse a number that has never been used
too_many_channels400The channel count has reached the limit (details carries max and current)Remove another channel first
channel_language_required400Exactly one transcription language was not specified (per_channel)Provide exactly 1 language in transcription_languages
channel_language_not_allowed400transcription_languages was provided in shared modeRemove the field
invalid_transcription_language400Invalid language codeConfirm the language code format is correct (such as zh-TW)
plan_feature_not_allowed403Exceeds the plan's limit on simultaneous recognition channels (details.field="max_stt_streams")Remove another channel or upgrade the plan
too_many_languages400The new channel's language would push the number of simultaneously recognized languages over the limit (plan cap; details carries max and received)Use an existing language or upgrade the plan
channel_action_while_paused400Channels cannot be added or disabled while pausedCall resume first
speaker_name_duplicate422speaker_name duplicates another channel's current nameUse a different name
invalid_parameter400speaker_name is too long (>100 characters) or contains control charactersFix the parameter per details.field
session_not_started400The recording has not startedCall start first
stt_start_failed500Speech recognition failed to start for this channel (the addition did not take effect)Retry later

Voice Translation - remove_channel (Disable a Channel / Multi-Channel)

Multi-channel mode only (v1.10.0)

Description

Disables a channel while a multi-channel recording is in progress. Once disabled, the channel no longer accepts new audio and no longer counts toward the recognition channel count; the billed channel count decreases immediately from the next minute, and the transcript and audio that channel has already produced are all preserved. Disabling leaves a wrap-up window of about 3 seconds so the channel's last sentence has a chance to make it into the transcript.

Notes:

  • The last remaining channel cannot be removed (returns channel_remove_not_allowed); use stop to end the recording
  • In shared mode, the first entry in channels[] carries the recognition for the whole session and cannot be removed (returns channel_remove_not_allowed)
  • A removed channel_id cannot be reused (see add_channel); to change a language, use set_channel_language — do not remove and re-add
  • Sending audio with that number after removal returns unknown_channel_id

Use Cases

  • A speaker leaves mid-session; release the channel to reduce billing
  • Reclaiming a temporarily opened guest microphone

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value remove_channel
channel_idintYesThe channel number to disable

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "remove_channel",
    "channel_id": 4
  }
}

Success Response

On success, a channel_status event is returned (reason: "removed", status: "removed"):

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 4,
      "speaker_name": "Director Lin",
      "transcription_languages": ["ja-JP"],
      "status": "removed"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "removed"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
not_multi_channel_session400This recording is not in multi-channel modeAvailable only with recognition_mode=multi_channel
channel_id_required400channel_id is missingProvide the number of the channel to disable
invalid_channel_id400channel_id is out of range (1–8)Use a valid channel number
unknown_channel_id400Unknown channel_id (it may have been removed already)Make sure the channel exists and has not been removed
channel_remove_not_allowed400The last remaining channel cannot be removed; or the first channel in shared modeUse stop to end the recording
channel_action_while_paused400Channels cannot be added or disabled while pausedCall resume first
session_not_started400The recording has not startedCall start first

Voice Translation - set_channel_language (Change a Channel's Language / Multi-Channel)

Multi-channel mode only (v1.10.0)

Description

Changes the transcription language of a single channel while a multi-channel recording is in progress. The channel_id and speaker identity (speaker_id) stay the same and the transcript remains continuous; the system switches that channel to the new language (it takes effect in about 4 seconds, during which the channel briefly produces no text and sends a channel_status event with reason: "language_change"). When the new language is one the session has not used yet, it is governed by the platform limit (10 simultaneously recognized languages) and the plan's limit on simultaneously recognized languages.

Do not substitute remove_channel + add_channel: channel numbers cannot be reused, and switching numbers would turn the same person into two different speakers in the transcript — with no way to fix it afterward.

Use Cases

  • The same speaker switches to another language mid-session
  • Correcting a channel language that was set incorrectly at start

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_channel_language
channel_idintYesTarget channel number
transcription_languagesstring[]ConditionalThe new transcription language (exactly 1; more than 1 returns invalid_parameter). Provide either this or language
languagestringConditionalThe new transcription language (single-value form). Provide either this or transcription_languages; providing both with different values returns invalid_parameter

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_channel_language",
    "channel_id": 2,
    "transcription_languages": ["ja-JP"]
  }
}

Success Response

On success, a channel_status event is returned (reason: "language_change", status: "preparing"); once the switch completes and the channel produces its first text, another event with status: "ready" arrives:

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 2,
      "speaker_name": "Alex",
      "transcription_languages": ["ja-JP"],
      "status": "preparing"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "language_change"
  }
}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
not_multi_channel_session400This recording is not in multi-channel modeAvailable only with recognition_mode=multi_channel
channel_id_required400channel_id is missingProvide the target channel number
invalid_channel_id400channel_id is out of range (1–8)Use a valid channel number
unknown_channel_id400Unknown channel_id (it may have been removed)Make sure the channel exists and has not been removed
channel_language_required400No new transcription language providedProvide transcription_languages[0] or language
invalid_parameter400Same as the current language, both fields provided with different values, or more than one languageFix the parameters according to details
invalid_transcription_language400Invalid language codeConfirm the language code format is correct (such as zh-TW)
channel_rebuild_too_frequent400Only one switch per channel is accepted within 5 seconds (details carries cooldown_seconds)Retry later
too_many_languages400The new language would push the number of simultaneously recognized languages over the platform limit (10) or the plan limitUse an existing language or upgrade the plan
channel_action_while_paused400The channel language cannot be changed while pausedCall resume first
session_not_started400The recording has not startedCall start first
channel_language_not_allowed400shared mode does not support changing a channel's languageThe session language is set at start and cannot be changed during the recording
stt_start_failed500Speech recognition failed to rebuild for this channel with the new languageRetry later (the channel waits for reconnection)

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "language_change"). The event carries channel_id to identify the channel.


Voice Translation - broadcast_go_live (Switch to the Live Phase)

Description

Switches from the broadcast standby phase (standby) to the live phase (live). After switching, STT/translation results begin broadcasting to viewers and start being written to the transcript.

Use Cases

  • The host confirms the equipment is working and starts the official broadcast
  • Switching from the warm-up phase to live streaming

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_go_live"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_phase_changed",
    "phase": "live",
    "message": "Broadcast started"
  }
}
FieldTypeDescription
phasestringThe new phase (live)
messagestringStatus description message

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
broadcast_not_enabled400Not broadcast modeConfirm type: "broadcast"
session_not_started400Speech recognition has not started, or the recording has ended (including while it is still being processed after ending)Call start first if it has not started

Note: If already in the live phase, a status message "Broadcast is already in progress" is returned and is not treated as an error.

If a sentence is in progress when this happens, it cannot be kept and is reported separately via segment_discarded (reason: "broadcast_go_live"). Standby-phase segments are never written to the final transcript anyway.


Voice Translation - broadcast_announcement (Send an Announcement)

Description

The host sends a custom message announcement to all viewers. Viewers receive an announcement event via SSE. The announcement message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.

Use Cases

  • Notifying viewers that the meeting is about to end
  • Sending an important reminder or announcement
  • One-way communication with viewers

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value broadcast_announcement
messagestringYesThe announcement message content

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_announcement",
    "message": "The meeting will end in 5 minutes"
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "Announcement sent"
  }
}

The SSE event viewers receive (with translations):

event: announcement
data: {"message":"The meeting will end in 5 minutes","translations":{"en-US":"The meeting will end in 5 minutes","ja-JP":"会議は5分後に終了します"}}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
broadcast_not_enabled400Not broadcast modeConfirm type: "broadcast"
invalid_parameter400Message is emptyProvide a valid message parameter

Voice Translation - set_standby_message (Set the Standby Phase Message)

Description

During the broadcast standby phase (standby), dynamically sets the message shown to viewers. This allows the host to enter standby mode and then set the waiting message, rather than being required to provide it at start.

The message is automatically translated into all translation languages, and the SSE event viewers receive includes a translations field.

Use Cases

  • After entering standby mode, dynamically set the waiting message shown to viewers
  • Update the text on the standby screen before going live
  • Reduce the required fields before starting the broadcast

Request Parameters

ParameterTypeRequiredDescription
actionstringYesFixed value set_standby_message
messagestringYesThe text displayed during the standby phase (translated for viewers of each language via the existing translation pipeline)

Request Example

{
  "type": "voice-translation",
  "data": {
    "action": "set_standby_message",
    "message": "The talk is about to begin, please wait..."
  }
}

Success Response

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "message": "Standby phase text updated"
  }
}

Event Viewers Receive

After a successful setting, all viewers in the standby phase receive an updated standby event via SSE:

event: standby
data: {"message":"The talk is about to begin, please wait...","translations":{"en-US":"The presentation is about to begin, please wait...","ja-JP":"プレゼンテーションがまもなく始まります。お待ちください..."}}

Note: The translations field contains the translation results for all translation languages. The frontend can display the corresponding translation based on the language the viewer selects.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Action
broadcast_not_enabled400Not broadcast modeConfirm type: "broadcast"
broadcast_not_in_standby400Not in the standby phaseCan be used only during the standby phase

Note: This action can be used only during the standby phase (standby). If the broadcast has already entered the live phase (live), an error is returned.


Response Events

The following are the commonly used WebSocket response events.

Note: This is not the complete list: translation_language_removed, summary_updated, summary_done, summary_error, upload_error and speakers_auto_merged are not documented on this page. See the event reference for all 36 events.

session_started - Session Started Successfully

After a start action succeeds, the server returns an event containing complete session initialization info. The frontend can distinguish the recording type via recording_type.

General recordings (transcribe / conversation / record):

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}

Broadcast mode (broadcast):

{
  "type": "voice-translation",
  "data": {
    "action": "session_started",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "broadcast",
    "recognition_mode": "multi_speaker",
    "phase": "standby",
    "viewer_count": 0,
    "queue_count": 0,
    "peak_viewers": 0,
    "total_viewers": 0,
    "resume_token": "L0VBAwIy... (43 characters)",
    "resume_grace_seconds": 45,
    "server_time": 1749550000000,
    "message": "Speech recognition started"
  }
}
FieldTypeDescription
session_idstringSession ID
task_idstringTask ID (can be used for subsequent API queries)
recording_typestringRecording type: transcribe, conversation, record, broadcast
recognition_modestringRecognition mode: single, multi_speaker
resume_tokenstringResume token for reconnection (43 characters). After a disconnect, reconnecting with this token within the resume_grace_seconds grace window rejoins the original session. It is sent ahead of time in session_started; keep it until the session ends.
resume_grace_secondsintReconnection grace period in seconds (default 45). This is wall-clock real time and keeps counting down after a disconnect.
server_timeint64Current server time in unix milliseconds. The client may use the pair (server_time, client time when received) to estimate clock skew for reference; the grace countdown is still based on the client wall-clock.
messagestringStatus description message
phasestringBroadcast phase: standby or live (broadcast mode only)
viewer_countintCurrent number of online viewers (broadcast mode only)
queue_countintNumber of viewers waiting in the queue (broadcast mode only)
peak_viewersintPeak number of viewers for this broadcast (broadcast mode only)
total_viewersintTotal cumulative number of viewers who have connected (broadcast mode only)
channel_modestringMulti-channel sub-mode: per_channel or shared (multi-channel mode only). This is also how the frontend confirms the server actually started in multi-channel mode
channelsarrayMulti-channel channel list (multi-channel mode only). Each entry carries channel_id, speaker_name, transcription_languages, status (preparing / ready / removed / error); transcription_languages is not present in shared mode

resume_ok - Resume Succeeded

The event returned by the server when a reconnection with resume_token succeeds within the resume_grace_seconds grace window after a disconnect. Upon receiving it, restart the audio stream just as you would after start (WebM/Opus must send a fresh container; PCM can continue directly), and align your local content using server_last_sid / server_last_offset_ms. If is_paused is true, the session was paused before the disconnect, so stay paused after resuming and send no audio. For the full resume handshake flow, see Connection and Authentication; for field details, see WebSocket Events Reference.

{
  "type": "voice-translation",
  "data": {
    "action": "resume_ok",
    "session_id": "550e8400-e29b-41d4-a716-446655440000",
    "task_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "recording_type": "transcribe",
    "recognition_mode": "single",
    "server_last_sid": 42,
    "server_last_offset_ms": 125000,
    "server_recording_ms": 127500,
    "is_paused": false,
    "settings": { ... },
    "message": "Session resumed"
  }
}
FieldTypeDescription
session_idstringSession ID (WS connection level; invalidated when the connection ends)
task_idstringTask ID (can be used for subsequent API queries)
recording_typestringRecording type: transcribe, conversation, record, broadcast
recognition_modestringRecognition mode: single, multi_speaker
server_last_sidintThe server's current last sentence number (sid). The client should ignore duplicate messages with sid ≤ this value.
server_last_offset_msint64The transcript timeline position (milliseconds) of the breakpoint, based on the length of audio already processed (not wall-clock; the audio timeline is frozen during the disconnect). The client uses it to place reconnected content at the correct timeline position.
server_recording_msint64The recording-head timestamp (milliseconds, including silence) that the transcript timeline resumes from after reconnect. The client uses it to align its recording-second header to the same timeline the transcript uses. The difference from server_last_offset_ms is the trailing audio (silence or not-yet-finalized speech) after the last finalized sentence before the disconnect. Omitted when 0.
is_pausedboolThe paused state the server considers authoritative after resuming: true = it was paused before the disconnect, so the client should stay paused (reopen the audio stream then pause immediately and send no audio); false = recording normally. Used to align the pause UI after a network reconnect or a full-page refresh. Omitted is treated as false (backward compatible with older servers).
settingsobjectThe recording settings currently held by the session (for state reconciliation, added in v1.6.4); see Connection and Authentication for details.
messagestringStatus description message (always "Session resumed")

Note: The two times must not be mixed: the transcript timeline (server_last_offset_ms) is based on the audio and freezes during a disconnect; the grace reconnection window (resume_grace_seconds) is wall-clock real time and keeps counting down during a disconnect. Deciding "whether reconnection is still possible" must use wall-clock, not the audio timeline position.

Multi-channel mode: the settings of resume_ok contains channel_mode and channels[] (each entry with channel_id, speaker_name, transcription_languages, status; transcription_languages is not present in shared mode), so that after resuming the frontend can confirm the server is still in multi-channel mode and that each channel's channel-to-language binding is unchanged. After the resume, each channel first shows preparing in the snapshot, and sends its own channel_status event (ready, reason: "reconnect") when it starts producing text. In shared mode, the other channels' status in the snapshot and events follows the first channel (see Shared Mode).


result - Recognition/Translation Result

Speech recognition and translation results. A single result event may contain origin (recognition result) and/or translations (translation results).

origin (speech recognition result):

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "origin": {
      "sid": 1,
      "language": "zh-TW",
      "text": "Hello, nice to meet you",
      "is_final": true,
      "speaker_id": "0",
      "detected_language": "zh-TW",
      "start_time": "00:05"
    }
  }
}
FieldTypeDescription
sidintSentence number, starting from 1
languagestringSource language code. In two-way mode and multi-language transcription this is the language determined by the system — in multi-language mode it is determined per sentence (so it varies within a single recording), and when the determination is unreliable the configured value is kept, so the field is never empty.
textstringThe recognized text
is_finalbooleanWhether it is the final result
speaker_idstringSpeaker ID. In multi-channel mode it is determined by the channel, format channel_{N} (such as channel_1)
speaker_labelstringSpeaker display label (optional). In multi-channel mode this is the speaker_name provided at start / add_channel (after a rename_speaker, the new label; equal to speaker_id when no name was provided)
detected_languagestringThe detected language. In two-way mode, this is determined automatically by the system.
start_timestringSentence start time (mm:ss); not sent during the broadcast standby phase; after going live, counts from 00:00.
channel_idintMulti-channel mode only: the source channel number of this sentence. Not present outside multi-channel mode

Multi-channel mode: each channel segments sentences independently, so sentences from different channels arrive interleaved and sid is not guaranteed to be sequential per channel; origin.language is the language bound to that channel.

translations (translation results):

{
  "type": "voice-translation",
  "data": {
    "action": "result",
    "translations": {
      "en-US": {
        "sid": 1,
        "text": "Hello, nice to meet you",
        "is_final": true
      }
    }
  }
}

Translation results are keyed by language code, and each language's translation object contains:

FieldTypeDescription
sidintSentence number
textstringThe translated text
is_finalbooleanWhether it is the final result
is_retranslationbooleanWhether it is a retranslation result (only during retranslate)

Multi-channel mode: translations does not carry channel_id; use sid to match back to origin for the source channel.


status - Generic Status Response

Used to confirm operations such as pause, resume, stop, set_name, tts_stop, and start_speaking.

{
  "type": "voice-translation",
  "data": {
    "action": "status",
    "status": "paused",
    "message": "Speech recognition paused"
  }
}
FieldTypeDescription
statusstringMachine-readable recording lifecycle state: live (resumed) / paused / ended (stopped). Only pause / resume / stop carry it; set_name etc. do not. ended is always sent before task_complete. Floating subtitle consumers should act on it: paused → freeze, ended → close the window, live → resume.
messagestringHuman-readable status text (format not guaranteed; do not parse — rely on the status field)

task_complete - Task Processing Complete

Triggered after stop when the audio file and transcript have been uploaded. task_id can be used to query task details via the REST API afterward.

  • It is always sent after status: "ended".
  • The transcript includes the summary, so this event waits until the title and summary have been generated. This usually takes a few seconds to tens of seconds, and longer for a long summary or a slow service, up to about 8 minutes.
  • The concurrent recording slot is released when this event is sent: you can start the next recording as soon as you receive it.
  • To detect completion, we recommend also supporting the Webhook or a REST query rather than relying on this event alone.
  • You also receive this event when a recording ends automatically after a long silence.
{
  "type": "voice-translation",
  "data": {
    "action": "task_complete",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "message": "Task processing complete"
  }
}

Recordings that received no audio at all: whether the recording ended with stop, insufficient credit, or automatically, if no audio was received during the whole recording you still receive status: "ended" and then this event, with noAudio: true. Such a recording has no transcript or audio file and its status is failed; a recording.failed Webhook is also sent (failure_source is no_audio). Do not try to read the transcript; tell the user that no sound was received instead. Minutes already elapsed are billed as usual.

{
  "type": "voice-translation",
  "data": {
    "action": "task_complete",
    "task_id": "550e8400-e29b-41d4-a716-446655440000",
    "noAudio": true,
    "message": "Task processing complete"
  }
}

A broadcast that ends during the standby phase (before going live) sends only status: "ended", not this event: no recording is created during the standby phase.

FieldTypeDescription
task_idstringRecording UUID, can be used for subsequent API queries
noAudiobooleanPresent only when no audio was received during the whole recording, and always true; absent when audio was recorded
messagestringStatus description

config_updated - Settings Update Complete

Triggered after the config action succeeds.

Note: Sent only when the configuration is accepted. A rejected configuration returns type: "error" instead, so your client must listen for that as well or the request appears to go unanswered.

{
  "type": "voice-translation",
  "data": {
    "action": "config_updated",
    "updated": ["terminology", "fuzzy_correction", "translation_dict"],
    "message": "Settings updated"
  }
}
FieldTypeDescription
updatedstring[]The setting types that were updated: terminology, fuzzy_correction, translation_dict
messagestringStatus message
terminology_effectivestring(Optional) Appears when terminology is updated during recording; a value of "next_turn" means the new terminology takes effect from the next sentence. Does not appear in the initial config
unknown_languagesstring[](Optional) Glossary language codes that could not be recognized (for example zh, chinese). Those entries will not take effect, but the config still succeeds
inactive_languagesstring[](Optional) Codes that are valid but are not used by this recording. Does not appear before the recording starts (the language list is not settled yet)
inactive_dict_languagesstring[](Optional) The same, for translation-dictionary target languages
homophone_conflictsobject[](Optional) Groups of terms in your glossary that share a pronunciation; each entry carries languages and terms. Not present before the recording has started (the language list is not settled yet)

Reading homophone_conflicts: when two terms that sound alike are registered together (公事包 and 公式包, say) and the transcript contains a third spelling with the same pronunciation, the system can only correct it to one of them, and which one is not guaranteed. Both terms themselves still work; the only ambiguity is which term an unregistered homophone misspelling is attributed to.

This is a warning, not an error — the glossary is still accepted and config still succeeds. If the pair matters to you, list the misspelling explicitly with fuzzy_correction so it is pinned to the term you want.

languages lists every language that shares the same term index. Chinese regional codes (zh-TW, zh-CN, zh-HK and so on) are treated as one group for glossary matching, so they share a single entry rather than each reporting one.

"homophone_conflicts": [
  { "languages": ["zh-TW"], "terms": ["公事包", "公式包"] }
]

tts_ready - TTS Audio Ready

TTS speech synthesis completion event. Contains the audio data and Word Boundary information (which can be used for a karaoke effect).

{
  "type": "voice-translation",
  "data": {
    "action": "tts_ready",
    "sid": 1,
    "language": "en-US",
    "transcript": "你好,很高興認識你",
    "text": "Hello, nice to meet you",
    "audio": "Base64EncodedMP3...",
    "format": "mp3",
    "duration_ms": 2500,
    "boundaries": [
      {"offset_ms": 0, "duration_ms": 350, "text_offset": 0, "word_length": 5, "text": "Hello"},
      {"offset_ms": 350, "duration_ms": 100, "text_offset": 5, "word_length": 1, "text": ","},
      {"offset_ms": 500, "duration_ms": 250, "text_offset": 7, "word_length": 4, "text": "nice"},
      {"offset_ms": 750, "duration_ms": 200, "text_offset": 12, "word_length": 2, "text": "to"},
      {"offset_ms": 950, "duration_ms": 350, "text_offset": 15, "word_length": 4, "text": "meet"},
      {"offset_ms": 1300, "duration_ms": 300, "text_offset": 20, "word_length": 3, "text": "you"}
    ]
  }
}
FieldTypeDescription
sidintSentence number
languagestringTTS language
transcriptstringThe original transcript (STT recognition result)
textstringThe translated text (TTS synthesis source)
audiostringBase64-encoded MP3 audio
formatstringAudio format (fixed value mp3)
duration_msintTotal audio duration (milliseconds)
boundariesarrayArray of Word Boundaries

Word Boundary Field Descriptions

FieldTypeDescription
offset_msintThe word's start time in the audio (milliseconds)
duration_msintThe word's duration (milliseconds)
text_offsetintPosition in the original string (character index)
word_lengthintWord length (number of characters)
textstringThe word content

tts_error - TTS Synthesis Failed

TTS synthesis failure event.

{
  "type": "voice-translation",
  "data": {
    "action": "tts_error",
    "sid": 1,
    "language": "en-US",
    "error": "translation_not_found",
    "message": "No translation available for language: en-US"
  }
}
FieldTypeDescription
sidintSentence number
languagestringTTS language
errorstringError code
messagestringError message
transcriptstringThe corresponding original transcript, to help the frontend locate the point of failure. This field is always present; it is an empty string "" when the sentence is not found

TTS Error Codes

Error CodeDescription
sentence_not_foundThe specified sentence was not found (the sid passed to tts_play does not exist)
translation_not_foundNo translation found for that language
tts_invalid_languageThe TTS language is not supported
tts_invalid_voiceThe voice name is invalid
tts_connection_failedCould not connect to the speech synthesis service
tts_timeoutSpeech synthesis timed out
tts_synthesis_failedSpeech synthesis failed

viewer_count - Viewer Count Update

Broadcast mode only

During a broadcast, the system monitors the viewer count and pushes this event to the host whenever it changes.

{
  "type": "voice-translation",
  "data": {
    "action": "viewer_count",
    "viewer_count": 45,
    "queue_count": 8,
    "peak_viewers": 50,
    "total_viewers": 123
  }
}
FieldTypeDescription
viewer_countintCurrent number of online viewers
queue_countintNumber of viewers waiting in the queue
peak_viewersintPeak number of viewers for this broadcast
total_viewersintTotal cumulative number of viewers who have connected

Note: This event is pushed only when the viewer count or queue count changes, to avoid unnecessary message traffic.


viewer_joined - Viewer Joined

Broadcast mode only

When a viewer joins the broadcast, the host receives this event.

{
  "type": "voice-translation",
  "data": {
    "action": "viewer_joined",
    "viewer_count": 5,
    "queue_count": 2
  }
}
FieldTypeDescription
viewer_countnumberCurrent number of viewers
queue_countnumberNumber waiting in the queue

viewer_left - Viewer Left

Broadcast mode only

When a viewer leaves the broadcast, the host receives this event.

{
  "type": "voice-translation",
  "data": {
    "action": "viewer_left",
    "viewer_count": 4,
    "queue_count": 1
  }
}
FieldTypeDescription
viewer_countnumberCurrent number of viewers
queue_countnumberNumber waiting in the queue

broadcast_phase_changed - Broadcast Phase Changed

Triggered when the broadcast phase switches from standby to live.

{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_phase_changed",
    "phase": "live",
    "message": "Broadcast started"
  }
}
FieldTypeDescription
phasestringThe new phase: standby or live
messagestringStatus description message

broadcast_recording_ready - Broadcast Recording Ready

Triggered after a broadcast goes live, returning the finalized task_id for this broadcast. The task_id in session_started is an initial value for the connection stage, not the final ID for this broadcast; you receive this event after going live regardless of whether the broadcast went through the standby phase first or started directly with broadcast_phase: "live" (the default). For subsequent operations (such as exchanging for a Floating Subtitle Feed Token), always use the task_id from this event — using the initial value returns no matching record.

  • When start goes live directly, this event always arrives after session_started.
  • When the host disconnects and, within the grace period, sends start again with the same broadcast_token (a takeover), this event on the new connection carries a new task_id. The broadcast therefore has two recordings, one before and one after the takeover, each completed and notified separately.
{
  "type": "voice-translation",
  "data": {
    "action": "broadcast_recording_ready",
    "task_id": "3f9a1c2e-..."
  }
}
FieldTypeDescription
task_idstringThe finalized recording ID for this broadcast (valid after going live)

speaker_renamed - Speaker Renamed

Speaker global rename completion event.

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_renamed",
    "speaker_id": "Guest-1",
    "new_label": "Manager Wang",
    "affected_sids": [1, 3, 5, 8]
  }
}
FieldTypeDescription
speaker_idstringThe resolved original speaker ID (even if the input was a display label, the event returns the original ID)
new_labelstringThe new display label
affected_sidsint[]The list of affected sentence numbers

speaker_reassigned - Speaker Identity Changed

Single-sentence speaker identity change completion event.

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_reassigned",
    "sid": 5,
    "old_speaker_id": "Guest-1",
    "new_speaker_id": "Guest-2",
    "new_speaker_label": "Lee Hsiao-hua"
  }
}
FieldTypeDescription
sidintThe changed sentence number
old_speaker_idstringThe original speaker ID
new_speaker_idstringThe new original speaker ID
new_speaker_labelstringThe new speaker display label (after applying speaker_aliases; equals new_speaker_id when no alias exists)

speakers_merged - Speakers Merged

Speaker merge completion event. After the merge, future recognition results for that source speaker are also automatically converted to the target speaker. This applies within the current recognition pass only: after a connection recovers, speakers are re-identified and you need to merge again.

{
  "type": "voice-translation",
  "data": {
    "action": "speakers_merged",
    "source_speaker_id": "Guest-2",
    "target_speaker_id": "Guest-1",
    "affected_sids": [3, 5, 7]
  }
}
FieldTypeDescription
source_speaker_idstringThe original ID of the merged speaker
target_speaker_idstringThe original ID of the merge target
affected_sidsnumber[]The list of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target)

To obtain the target speaker's display label, query speaker_aliases or the next init_metadata event.


language_switch_start - Language Switch Started

Language switch start event, sent after the switch_language action is triggered.

{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_start",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "total_segments": 15
  }
}
FieldTypeDescription
translation_languagestringThe single language for this operation (the new target for a replace, or the language added by op:add)
translation_languagesstring[]Authoritative snapshot of the current full set of translation languages. For a multi-language op:add this is the complete set including existing languages. Clients should overwrite their local language set with this directly, rather than inferring "append" vs "replace" from translation_language (when only 1 language exists, an op:add expanding to a 2nd would be misread as a replace and drop the existing language).
total_segmentsintThe number of sentences that need retranslation

batch_retranslation - Batch Retranslation Result

Batch retranslation result event, sent sentence by sentence during the language switch process.

{
  "type": "voice-translation",
  "data": {
    "action": "batch_retranslation",
    "sid": 3,
    "translations": {
      "ja-JP": {
        "sid": 3,
        "text": "今日はプロジェクトの進捗について話し合いましょう",
        "is_final": true,
        "is_retranslation": true
      }
    }
  }
}
FieldTypeDescription
sidintSentence number
translationsobjectTranslation results (same format as result's translations)

language_switch_done - Language Switch Complete

Language switch completion event.

{
  "type": "voice-translation",
  "data": {
    "action": "language_switch_done",
    "translation_language": "ja-JP",
    "translation_languages": ["en-US", "ja-JP"],
    "success_count": 15,
    "failed_count": 0
  }
}
FieldTypeDescription
translation_languagestringThe single language for this operation
translation_languagesstring[]Authoritative snapshot of the current full translation-language set (same as language_switch_start; overwrite the local set).
success_countintThe number of successfully translated sentences
failed_countintThe number of sentences that failed to translate

tts_mode_changed - TTS Mode Changed

TTS playback mode change event.

{
  "type": "voice-translation",
  "data": {
    "action": "tts_mode_changed",
    "tts_mode": "async"
  }
}
FieldTypeDescription
tts_modestringThe new mode: sync or async

language_switched - Two-Way Language Switch Complete

Two-way mode (conversation) language switch completion event. Triggered after switch_language successfully switches the STT source language in two-way mode.

{
  "type": "voice-translation",
  "data": {
    "action": "language_switched",
    "language": "en-US",
    "translation_language": "zh-TW",
    "message": "Language switched"
  }
}
FieldTypeDescription
languagestringThe new active language (STT source)
translation_languagestringThe new translation target language
messagestringStatus message

tts_updated - Two-Way TTS Settings Updated

Two-way mode (conversation) TTS settings update event. Triggered after set_tts successfully updates the TTS toggle or voice settings.

{
  "type": "voice-translation",
  "data": {
    "action": "tts_updated",
    "tts_enabled": true,
    "tts_config": {
      "zh-TW": { "voice": "zh-TW-HsiaoChenNeural", "speaking_rate": 1.0 },
      "en-US": { "voice": "en-US-GuyNeural", "speaking_rate": 1.2 }
    }
  }
}
FieldTypeDescription
tts_enabledbooleanWhether TTS is enabled
tts_configobjectThe TTS settings for each language (voice, speaking_rate)

conversation_mode_changed - Conversation Mode Changed

Two-way mode (conversation) conversation mode change event. Triggered after switch_conversation_mode successfully switches between auto/manual mode.

{
  "type": "voice-translation",
  "data": {
    "action": "conversation_mode_changed",
    "conversation_mode": "manual"
  }
}
FieldTypeDescription
conversation_modestringThe new conversation mode: auto or manual

speaker_language_changed - User Language Changed

Two-way mode (conversation) user language change event. Triggered after set_speaker_language successfully changes a user's language, including the complete language mapping after the change.

{
  "type": "voice-translation",
  "data": {
    "action": "speaker_language_changed",
    "speaker_language_map": {
      "1": "ja-JP",
      "2": "en-US"
    }
  }
}
FieldTypeDescription
speaker_language_mapobjectThe user language mapping after the change (keys are user number strings)

speaking_speed_changed - Speaking Speed Changed

Speaking speed change event during recording. Triggered after set_speaking_speed successfully applies the new speed (STT rebuild complete), returning the applied speed level.

{
  "type": "voice-translation",
  "data": {
    "action": "speaking_speed_changed",
    "speaking_speed": "slow"
  }
}
FieldTypeDescription
speaking_speedstringThe applied speed: very_slow / slow / normal / fast / very_fast

channel_status - Channel Status Changed

Multi-channel mode only (v1.10.0)

In multi-channel mode (recognition_mode: "multi_channel"), fired whenever a channel's status changes: the success responses of add_channel / remove_channel, setting changes from set_channel_language and set_speaking_speed, pause/resume and automatic reconnection after a disconnect, and channel failures.

{
  "type": "voice-translation",
  "data": {
    "action": "channel_status",
    "channel": {
      "channel_id": 3,
      "speaker_name": "Lee",
      "transcription_languages": ["zh-TW"],
      "status": "preparing"
    },
    "active_channels": 3,
    "stt_stream_count": 3,
    "reason": "added"
  }
}
FieldTypeDescription
channelobjectThe channel whose status changed, with channel_id, speaker_name, transcription_languages, status
active_channelsintCurrent number of active channels (number of speakers)
stt_stream_countintCurrent number of recognition channels being counted (billed on this count as each minute begins); always 1 in shared mode
reasonstringOptional. Why this status change happened; see the table below

status state machine

Each time a channel starts or prepares recognition again, it first enters preparing and transitions to ready once it produces its first text:

statusDescription
preparingPreparing (creating the channel or applying a new setting). Speech on this channel appears a few seconds late, but is not lost; the one exception is a sentence in progress when the setting is applied, which cannot be kept (see segment_discarded)
readyThe channel has received its first recognition result
removedDisabled by remove_channel
errorSpeech recognition on this channel failed and cannot recover automatically

Shared mode: the other channels' preparing / ready / error follow the first channel, and each channel receives its own event when the first channel's status changes; removed follows each channel's own status. See Shared Mode.

reason values

reasonDescription
addedAdded by add_channel
language_changeLanguage switched by set_channel_language
reconnectAutomatic reconnection after the channel failed unexpectedly
resumedResumed after a pause
removedDisabled by remove_channel
speaking_speedNew setting applied channel by channel via set_speaking_speed
stt_errorSpeech recognition failure

segment_discarded - Segment Discarded

Tells you that a sid has been discarded and will receive no further events — no is_final: true original text, and no translation for that segment.

Some operations interrupt recognition. If a sentence happens to be in progress at that moment, it cannot be kept. Your client may already have received interim results (is_final: false) for it, so on receiving this event, clear that sid from any "translating" state.

Automatic conversation mode never receives this event — in that mode, a sentence in progress is finalized and sent with is_final: true, and its content is kept in the transcript.

{
  "type": "voice-translation",
  "data": {
    "action": "segment_discarded",
    "sid": 7,
    "reason": "speaking_speed",
    "channel_id": 1
  }
}
FieldTypeDescription
sidintThe discarded segment number
reasonstringspeaking_speed (set_speaking_speed) / language_change (set_channel_language) / reconnect (automatic reconnection after the recognition connection failed) / resumed (session resume or resuming after a pause) / broadcast_go_live (moving from standby to live)
channel_idintOptional. Source channel number; sent only in multi-channel mode

The first four reason values come from the same set as channel_status, so you can match both events back to the same operation.

You may already have received this event even when set_speaking_speed fails (set_speaking_speed_failed) — recognition is interrupted before the rebuild, so the segment cannot be kept whether the rebuild succeeds or not.


segment_uploaded - Audio Segment Upload Complete

Audio segment upload completion event. Triggered each time an audio segment is successfully uploaded to cloud storage; can be used to show upload progress on the frontend.

{
  "type": "voice-translation",
  "data": {
    "action": "segment_uploaded",
    "segment_index": 0,
    "duration_sec": 30.5
  }
}
FieldTypeDescription
segment_indexnumberSegment index (starting from 0)
duration_secnumberThe duration of this segment (seconds)

stt_event - STT Connection Status Event

STT connection status event. Triggered when the connection status of the speech recognition service changes; can be used to show the STT service status on the frontend.

{
  "type": "voice-translation",
  "data": {
    "action": "stt_event",
    "event": "reconnected",
    "message": "STT reconnected"
  }
}
FieldTypeDescription
eventstringEvent type: session_started (recognition session established), session_stopped (recognition session ended), canceled (recognition aborted; see message for the reason), reconnecting (connection lost, reconnecting automatically), reconnected (reconnected)
messagestringEvent description message. The wording is not guaranteed to be stable, so always branch on event

error - Error Event

Triggered when an operation fails or a system anomaly occurs.

{
  "type": "error",
  "data": {
    "error_code": "session_not_started",
    "severity": "error",
    "message": "Session not started",
    "context": "voice-translation",
    "request_id": "req_abc123xyz789",
    "timestamp": "2026-01-15T10:30:45.123Z"
  }
}
FieldTypeDescription
error_codestringError code (for programmatic handling)
severitystringSeverity: fatal / error / warning
messagestringHuman-readable error message
contextstringError source category
request_idstringRequest tracking ID
timestampstringTime the error occurred (ISO 8601)

Severity Descriptions

severityDescriptionRecommended Action
fatalFatal errorStop the service and require reconnection
errorOperation failedShow an error notice and allow retry
warningWarningShow a warning without blocking the operation

A warning-level error does not mean the recording has ended; examples are stt_silence_warning (about to end for lack of speech) and broadcast_standby_warning (the standby phase is about to reach its time limit). The recording continues; when it actually ends you receive the corresponding fatal error and status: "ended".

For the full list of error codes, refer to Error Code Reference.


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026