API Docs

SSE API

Note: The URL used in this document (vas-poc.vurbo.ai) is the planned deployment URL. The official URL will be announced separately after launch.


Table of Contents


Connection Information

ItemValue
Base Pathhttps://vas-poc.vurbo.ai/api/v1/sse
ProtocolHTTP + Server-Sent Events (SSE)
Data Formattext/event-stream
AuthenticationHeader X-API-Key: {KEY}

Authentication

SSE APIs that require authentication accept two delivery methods for the API key (both are supported):

# Method A: HTTP Header (recommended, better security)
X-API-Key: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

# Method B: Query string (native browser EventSource fallback)
?api_key=vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

Note: The native browser EventSource API does not support custom headers. You can use the ?api_key= query string instead, or use the fetch API with a ReadableStream / an SSE client library that supports headers. In query string mode, the API key appears in the URL, so avoid writing the full URL into server logs or leaking it in screenshots.


Floating Subtitle SSE

A read-only "floating subtitle" transcript stream for an in-progress recording (source-language original text + target-language translations), subscribed over an independent connection with cross-device/window support. The base path is https://vas-poc.vurbo.ai (the real-time service). First exchange for a feed_token via POST /api/v1/auth/tasks/{taskId}/subtitle-feed-token, then connect to GET /tasks/{task_id}/subtitle?feed_token=....

The owner can also enable sharing so that other on-site audience members connect with a share secret (read-only, not charged separately; the audience limit is server-configured, default 10, excluding the owner). While connected, the viewers event reports the current viewer count.

See Floating Subtitle SSE spec for details.


Broadcast SSE API

The Broadcast SSE API provides a live subtitle streaming feature, allowing viewers to watch real-time transcription and translation content through a share link.

Note: The base path for Broadcast SSE is https://vas-poc.vurbo.ai/broadcast, which differs from the other SSE APIs.

GET /broadcast/{token}/text (Viewer Live Subtitle Stream)

Description

Viewers connect using a share token to receive an SSE stream of real-time transcription and translation.

Use Cases

  • Viewers watching live subtitles
  • Multilingual translation subtitle display
  • TTS audio playback

Authentication

Token authentication (no API key required): verified through the {token} in the URL path.

Request Parameters

ParameterTypeRequiredDescription
tokenstringYesBroadcast share token (4-character short code, a-z0-9, path parameter)
langstringNoFilter for a specific translation language (e.g., en-US)
ttsbooleanNoWhether to enable TTS (true / false, default false)
viewer_access_tokenstringConditionalViewer access token (required for password-protected broadcasts)

Password Protection Note: When a broadcast is set to password-protected, viewers must first obtain a viewer_access_token through the password verification API, then include this token in the query parameters of the SSE connection.

Request Example

// Receive all languages
const eventSource = new EventSource(
  'https://vas-poc.vurbo.ai/broadcast/a3f9/text'
);

// Receive only English translation
const eventSource = new EventSource(
  'https://vas-poc.vurbo.ai/broadcast/a3f9/text?lang=en-US'
);

// Receive English translation and enable TTS
const eventSource = new EventSource(
  'https://vas-poc.vurbo.ai/broadcast/a3f9/text?lang=en-US&tts=true'
);

Event Types

EventDescriptionNotes
connectedConnection confirmation-
queuedAdded to the waiting queueQueueing mechanism
admittedEntered live from the queueQueueing mechanism
originOriginal text (STT)-
translationTranslation result-
tts_readyTTS audio ready-
pausedBroadcast pausedHost paused manually (a host disconnection is reported as host_disconnected)
resumedBroadcast resumedHost resumed
endedBroadcast ended-
kickedRemovedViewer management
errorError-
speaker_renamedSpeaker renamed-
speaker_reassignedSingle-sentence speaker change-
speakers_mergedSpeakers merged-
recording_startedNew recording startedViewers should clear the previous subtitles on screen
max_viewers_changedViewer limit changed-
host_disconnectedHost temporarily offline and reconnectingThe broadcast is frozen, not ended
host_reconnectedHost has reconnecteddata.resumed indicates whether playback resumed as well
standbyStandby phase notification-
phase_changedPhase change notification-
announcementHost announcement-

Event Format

connected:

{
  "available_langs": ["en-US", "ja-JP"],
  "tts_languages": ["en-US"],
  "phase": "standby",
  "recognition_mode": "single"
}
FieldTypeDescription
available_langsarrayList of available translation languages
tts_languagesarrayList of languages with TTS enabled (an empty array means no TTS)
phasestringBroadcast phase: standby (preparing) or live (active)
recognition_modestringRecognition mode: single (single speaker) or multi_speaker (multi-speaker diarization)

The connected event does not carry the list of transcription languages. Retrieve the channel's public information (name, transcription_languages, translation_languages, TTS voices) with GET /api/v1/viewer/broadcasts/{token} instead - see Viewer API.

queued:

{
  "position": 3,
  "estimated_wait": "約 2 分鐘"
}
FieldTypeDescription
positionnumberPosition in the queue (1 = next up)
estimated_waitstringEstimated wait time. Server-fixed value in Traditional Chinese: 少於 1 分鐘 ("less than 1 minute") or 約 N 分鐘 ("about N minutes")

admitted:

{
  "message": "已進入直播"
}

Several broadcast events carry a message string that the server returns as a fixed Traditional Chinese value (已進入直播, 廣播已暫停, 廣播已恢復, 廣播已結束, 廣播已開始). These are display strings only - render them as-is or substitute your own copy, and never parse or compare against them.

origin:

{
  "sid": 1,
  "text": "Hello everyone",
  "speaker_id": "Guest-1",
  "speaker_label": "Guest-1",
  "start_time": "00:05",
  "is_final": true
}
FieldTypeDescription
sidnumberSentence ID
textstringOriginal text content
speaker_idstringOptional. Original speaker ID (immutable); sent only in multi-speaker diarization mode, omitted in single-speaker mode
speaker_labelstringDisplay label (after applying speaker_aliases; equals speaker_id when no alias exists)
start_timestringStart time (mm:ss), counted from 00:00 once the broadcast goes live. Standby content is never sent to viewers
is_finalbooleanWhether this is the final result

translation:

{
  "sid": 1,
  "language": "en-US",
  "text": "Hello everyone",
  "speaker_id": "Guest-1",
  "speaker_label": "Royx",
  "is_final": true
}
FieldTypeDescription
sidnumberThe corresponding sentence ID
languagestringTranslation language
textstringTranslated content
speaker_idstringOriginal speaker ID (multi-speaker conversation mode; immutable)
speaker_labelstringDisplay label (after applying speaker_aliases)
is_finalbooleanWhether this is the final result

tts_ready:

{
  "sid": 1,
  "language": "en-US",
  "transcript": "Hello, hi everyone",
  "text": "Hello everyone",
  "audio": "//uQxAAAAAANIAAAAAExBTUUzLjEwMFVVVV...",
  "format": "mp3",
  "duration_ms": 2340,
  "boundaries": [
    {"offset_ms": 0, "duration_ms": 320, "text": "Hello", "text_offset": 0, "word_length": 5},
    {"offset_ms": 320, "duration_ms": 280, "text": "everyone", "text_offset": 6, "word_length": 8}
  ]
}
FieldTypeDescription
sidnumberThe corresponding sentence ID
languagestringTTS language
transcriptstringOriginal transcript (source text)
textstringTranslated text
audiostringBase64-encoded MP3 audio
formatstringAudio format, fixed as "mp3"
duration_msnumberAudio duration (milliseconds)
boundariesarrayWord boundaries (optional, see table below)

Word boundary fields (each object in the boundaries array):

FieldTypeDescription
offset_msnumberThe word's start time in the audio (ms)
duration_msnumberThe word's pronunciation duration (ms)
textstringThe word text
text_offsetnumberThe word's starting position in the text
word_lengthnumberThe word's character length

Note:

  • The host must specify which languages enable TTS via the tts_config parameter in the start command
  • Only viewers who subscribed to that language and enabled TTS will receive this event
  • It is sent only during the live phase; no TTS is sent during the standby phase

paused:

{
  "message": "廣播已暫停"
}
FieldTypeDescription
messagestringNotification message

A host going offline and coming back are separate events (host_disconnected / host_reconnected), not reason values on paused.

resumed:

{
  "message": "廣播已恢復"
}

ended:

{
  "message": "廣播已結束"
}
FieldTypeDescription
messagestringNotification message

All three events carry only message — no end reason and no timestamps. The message text is for display only; determine state from the event name itself, not from this string.

kicked:

{
  "message": "Kicked by host"
}

message is for display only, never parse it. When the host removes a viewer the value is the fixed string Kicked by host; when the removal is caused by an access-type or passcode change, the value is a Traditional Chinese explanation of that change.

error:

{
  "error_code": "broadcast_session_ended",
  "severity": "error",
  "message": "Broadcast session ended",
  "context": "broadcast",
  "request_id": "req_abc123xyz789",
  "timestamp": "2025-12-05T10:30:45.123Z"
}

Sentence-level errors (such as a translation failure for a specific language) additionally carry sid and translation_language, making it easy for the frontend to flag which language failed for a given sentence:

{
  "error_code": "llm_content_filtered",
  "severity": "warning",
  "message": "Content filtered",
  "context": "translation",
  "sid": 5,
  "translation_language": "ja-JP",
  "timestamp": "2026-04-26T10:30:45.123Z"
}
FieldTypeDescription
error_codestringError code
severitystringSeverity: warning / error / fatal
messagestringError message (a fixed English display string returned by the server; display only, never parse it)
contextstringThe context in which the error occurred (e.g., broadcast, translation)
sidintOptional. The sentence number for a sentence-level error (e.g., when that sentence's translation fails)
translation_languagestringOptional. The target language that failed to translate (viewers can use this to determine whether a specific language failed for that sentence)
request_idstringOptional. Request tracking ID. Present only on connection-phase errors (non-200 HTTP status); sentence-level and session-level errors omit it
timestampstringTime the error occurred (ISO 8601)

speaker_renamed:

Multi-speaker conversation mode only. Sent when the host performs a global speaker rename.

{
  "speaker_id": "Guest-1",
  "new_label": "Royx",
  "affected_sids": [1, 3, 5, 7]
}
FieldTypeDescription
speaker_idstringThe resolved original speaker ID (even if the input is a display label, the event returns the original ID)
new_labelstringNew display label (e.g., Royx)
affected_sidsarrayList of affected sentence IDs

speaker_reassigned:

Multi-speaker conversation mode only. Sent when the host changes the speaker of a single sentence.

{
  "sid": 3,
  "old_speaker_id": "Guest-1",
  "new_speaker_id": "Guest-2",
  "new_speaker_label": "Amy"
}
FieldTypeDescription
sidnumberThe sentence ID that was modified
old_speaker_idstringOriginal speaker ID (e.g., Guest-1)
new_speaker_idstringThe new original speaker ID (e.g., Guest-2)
new_speaker_labelstringNew speaker display label (after applying speaker_aliases; equals the original ID when no alias exists)

speakers_merged:

Multi-speaker conversation mode only. Sent when the host merges speakers. After merging, all sentences belonging to that speaker are reassigned to the target speaker.

{
  "source_speaker_id": "Guest-2",
  "target_speaker_id": "Guest-1",
  "target_speaker_label": "Manager Wang",
  "affected_sids": [3, 5, 7]
}
FieldTypeDescription
source_speaker_idstringThe original speaker ID being merged (e.g., Guest-2)
target_speaker_idstringThe original speaker ID of the merge target (e.g., Guest-1)
target_speaker_labelstringTarget speaker display label (after applying speaker_aliases; equals the original ID when no alias exists)
affected_sidsarrayList of affected sentence IDs: sentences that belonged to the source speaker, plus the target speaker's existing sentences whose display name changed because of the merge (for example, when the source speaker's custom name is carried over to the target)

standby:

When a viewer connects during the standby phase, this event is received immediately after the connected event, indicating that the broadcast has not yet officially started. The host can dynamically update the standby message via the WebSocket set_standby_message action; after the update, all viewers receive a new standby event.

{
  "message": "The presentation is about to begin, please wait...",
  "translations": {
    "en-US": "The presentation is about to begin, please wait...",
    "ja-JP": "プレゼンテーションがまもなく始まります。お待ちください..."
  }
}
FieldTypeDescription
messagestringThe message displayed during the standby phase (original text)
translationsobjectTranslation results (optional); the key is the language code and the value is the translated text

phase_changed:

Sent when the broadcast switches from the standby phase to the active phase.

{
  "phase": "live",
  "message": "Broadcast has started"
}
FieldTypeDescription
phasestringThe new phase: live (active phase)
messagestringPhase change message

announcement:

An announcement message sent by the host; all viewers receive it.

{
  "message": "The meeting will end in 5 minutes",
  "translations": {
    "en-US": "The meeting will end in 5 minutes",
    "ja-JP": "会議は5分後に終了します"
  }
}
FieldTypeDescription
messagestringThe announcement content (original text)
translationsobjectTranslation results (optional); the key is the language code and the value is the translated text

Heartbeat Mechanism

The SSE connection uses a heartbeat to keep the connection alive:

  • Interval: 15 seconds
  • Format: SSE comment (starting with :)
  • The frontend does not need to handle it; the browser automatically ignores it
: heartbeat

Error Responses

Error CodeHTTP StatusDescriptionRecommended Handling
broadcast_session_not_found404Broadcast not foundConfirm the token is correct
broadcast_session_ended410Broadcast endedNotify the user that the broadcast has ended
broadcast_capacity_exceeded503Viewer capacity reachedJoin the waiting queue

Note: If an SSE endpoint encounters an unexpected internal exception, it may return internal_error (consistent with how WebSocket handles a single-message failure); expected domain errors return the corresponding error code (e.g., sse_translation_failed).

Frontend Example

function connectBroadcast(token, lang = null) {
  let url = `https://vas-poc.vurbo.ai/broadcast/${token}/text`;
  if (lang) {
    url += `?lang=${lang}`;
  }

  const eventSource = new EventSource(url);

  eventSource.addEventListener('connected', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Connected, phase: ${data.phase}, recognition mode: ${data.recognition_mode}`);
    console.log(`Available translations: ${data.available_langs.join(', ')}`);
  });

  eventSource.addEventListener('queued', (e) => {
    const data = JSON.parse(e.data);
    console.log(`In queue, position: ${data.position}, estimated wait: ${data.estimated_wait}`);
  });

  eventSource.addEventListener('admitted', (e) => {
    console.log('Entered live');
  });

  eventSource.addEventListener('origin', (e) => {
    const data = JSON.parse(e.data);
    console.log(`[${data.start_time}] ${data.text}`);
  });

  eventSource.addEventListener('translation', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Translation (${data.language}): ${data.text}`);
  });

  eventSource.addEventListener('tts_ready', (e) => {
    const data = JSON.parse(e.data);
    // Decode the Base64 audio and play it
    const byteCharacters = atob(data.audio);
    const byteNumbers = new Array(byteCharacters.length);
    for (let i = 0; i < byteCharacters.length; i++) {
      byteNumbers[i] = byteCharacters.charCodeAt(i);
    }
    const blob = new Blob([new Uint8Array(byteNumbers)], { type: 'audio/mpeg' });
    const audio = new Audio(URL.createObjectURL(blob));
    audio.play();
  });

  eventSource.addEventListener('paused', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Broadcast paused: ${data.message}`);
  });

  eventSource.addEventListener('resumed', (e) => {
    console.log('Broadcast resumed');
  });

  eventSource.addEventListener('ended', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Broadcast ended: ${data.message}`);
    eventSource.close();
  });

  eventSource.addEventListener('kicked', (e) => {
    console.log('You have been removed');
    eventSource.close();
  });

  eventSource.addEventListener('speaker_renamed', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Speaker renamed: ${data.speaker_id} → ${data.new_label}`);
    console.log(`Affected sentences: ${data.affected_sids.join(', ')}`);
    // Update the speaker display name for all affected sentences
  });

  eventSource.addEventListener('speaker_reassigned', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Speaker of sentence ${data.sid} changed from ${data.old_speaker_id} to: ${data.new_speaker_label}`);
    // Update the speaker display name for that sentence
  });

  eventSource.addEventListener('standby', (e) => {
    const data = JSON.parse(e.data);
    // Display the translation matching the viewer's selected language
    const displayLang = 'en-US'; // The language the viewer selected
    const displayMessage = data.translations?.[displayLang] || data.message;
    console.log(`Standby phase: ${displayMessage}`);
    // Show the waiting screen
  });

  eventSource.addEventListener('phase_changed', (e) => {
    const data = JSON.parse(e.data);
    console.log(`Phase changed: ${data.phase} - ${data.message}`);
    // Remove the waiting screen and start displaying subtitles
  });

  eventSource.addEventListener('announcement', (e) => {
    const data = JSON.parse(e.data);
    // Display the translation matching the viewer's selected language
    const displayLang = 'en-US'; // The language the viewer selected
    const displayMessage = data.translations?.[displayLang] || data.message;
    console.log(`Announcement: ${displayMessage}`);
    // Show the announcement message
  });

  eventSource.addEventListener('error', (e) => {
    if (e.data) {
      const error = JSON.parse(e.data);
      console.error(`Error [${error.error_code}]: ${error.message}`);
    }
    eventSource.close();
  });

  return eventSource;
}

REST API also available: For the channel's public information (name, language lists, TTS voices), see Viewer API.


GET /api/v1/sse/history/transcribe/{taskId} (Retrieve Conversation History)

Description

Loads the complete conversation history for the specified task, including all sentences and the summary. Delivered one item at a time via an SSE stream.

Use Cases

  • Viewing the recording details page
  • Loading the historical transcript

Authentication

Header: X-API-Key: YOUR_API_KEY

Request Parameters

ParameterTypeRequiredDescription
taskIdstringYesRecording ID (path parameter)

Request Example

// Use the fetch API (because EventSource does not support headers)
async function connectSSE(taskId, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/${taskId}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );
  const reader = response.body.getReader();
  // ... handle SSE events
}

Event Sequence

1. connected        → Connection confirmation
2. init_metadata    → Send task metadata
3. init_sentence    → Send sentences one at a time (repeated N times)
4. init_summary     → Send the summary
5. init_done        → Initialization complete

Event Format

connected:

{"message": "History service connected (taskId: xxx)"}

init_metadata:

{
  "task_id": "550e8400-e29b-41d4-a716-446655440000",
  "title": "Meeting Notes",
  "created_at": "2025-12-17T10:00:00Z",
  "type": "transcribe",
  "has_speaker_diarization": true,
  "transcription_languages": ["zh-TW"],
  "translation_languages": ["en-US"],
  "summary_template": "general",
  "summary_language": "zh-TW",
  "speaker_aliases": {"speaker_1": "Manager Wang"}
}

speaker_aliases is a mapping of "original speaker ID → display name"; it is {} (an empty object, not an array) when there are no aliases. The frontend can use this mapping to run a duplicate-name pre-check before renaming a speaker (added in v1.3.12).

init_sentence:

{
  "sid": 1,
  "origin": "Hello, nice to meet you",
  "translations": {
    "en-US": "Hello, nice to meet you"
  },
  "start_time": "00:05",
  "speaker_id": "speaker_1",
  "speaker_label": "Manager Wang"
}

If a sentence has a translation failure, it additionally carries a translation_errors field (only present when there is a failure), so the frontend can distinguish between "that language was not scheduled for translation" (the key is missing from translations) and "translated but failed" (the key is present in translation_errors). The same language may have both an older translation and a failure record (a failed retranslation keeps the previous translation), so read both fields together:

{
  "sid": 5,
  "origin": "Sentence with sensitive words",
  "translations": {
    "en-US": "Sensitive sentence"
  },
  "translation_errors": {
    "ja-JP": "llm_content_filtered"
  },
  "start_time": "00:25",
  "speaker_id": "speaker_1",
  "speaker_label": "Manager Wang"
}
FieldTypeDescription
sidintSentence number
originstringOriginal text
translationsobjectTranslation results (optional); the key is the language code and the value is the translated text
translation_errorsobjectOptional. Translation failure error codes; the key is the language code and the value is the error_code (e.g., llm_content_filtered)
channel_idnumberOptional. (Multi-channel only) which physical channel this sentence came from; present only for multi-channel recordings, omitted for single-channel
start_timestringStart time (mm:ss format)
speaker_idstring|nullOriginal speaker ID (immutable, e.g., speaker_1); the source for target_speaker_id in PATCH /speakers/reassign (flipped in v1.5.3: previously the display name)
speaker_labelstring|nullDisplay label (the human-readable name after applying speaker_aliases, e.g., Manager Wang); equals speaker_id when no alias exists (added in v1.5.3 to replace the original speaker_id display semantics)

init_summary:

{
  "text": "This is a summary of the meeting notes...",
  "mode": "builtin",
  "template": "meeting",
  "plain_text": false,
  "summary_language": "zh-TW"
}

mode, template, plain_text and summary_language are always present (summary_language is the language of this summary and is null when there is no summary); custom mode also carries prompt_snapshot, and an automatic downgrade adds fallback_level / dropped_segments. See the history endpoint for the full field list.

init_done:

{"totalSentences": 10}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Handling
recording_not_found404Recording not foundConfirm the taskId is correct
sse_transcript_not_found404Transcript not foundThe recording may not have finished processing yet

Frontend Example

// Use the fetch API to handle SSE (you must parse the event-stream yourself)
async function loadHistory(taskId, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/${taskId}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    const text = decoder.decode(value);
    // Parse the SSE format: event: xxx\ndata: {...}\n\n
    const events = parseSSE(text);

    for (const event of events) {
      if (event.type === 'init_metadata') {
        console.log('Task info:', event.data.title);
      } else if (event.type === 'init_sentence') {
        console.log(`[${event.data.start_time}] ${event.data.origin}`);
        if (event.data.translations) {
          console.log(`Translation: ${event.data.translations['en-US']}`);
        }
      } else if (event.type === 'init_done') {
        console.log('Loading complete');
      }
    }
  }
}

GET /api/v1/sse/retranslate/{taskId} (Retranslate Full Transcript)

Description

Retranslates all sentences of the specified task into the target language. Translation results are delivered one at a time via an SSE stream.

Use Cases

  • Switching the display language
  • Updating the translation content

Authentication

Header: X-API-Key: YOUR_API_KEY

Request Parameters

ParameterTypeRequiredDescription
taskIdstringYesRecording ID (path parameter)
targetLangstringYesTarget language code
segmentedstringNo1 declares that the client supports segmented retranslation (v1.18.0)
fromSidnumberNoOnly together with segmented=1: start from this sentence (inclusive) (v1.18.0)
expectedRevisionnumberNoTranscript revision; on a mismatch nothing is translated or charged and transcript_revision_conflict is returned (v1.18.0)

Use segmented retranslation for long transcripts: without segmented=1, the whole transcript is translated at once; if it is too long to translate in one request, HTTP 422 retranslate_segmentation_required is returned before the stream starts (no charge). With segmented=1, each request translates one segment; done carries revision, and when the segment was cut short also truncated: true and nextSid. Send the next segment with fromSid={nextSid} until done no longer carries truncated. Each segment is saved and charged separately. See reference/sse/retranslate.md for the full rules.

Request Example

// Use the fetch API (because EventSource does not support headers)
async function retranslateSSE(taskId, targetLang, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/retranslate/${taskId}?targetLang=${targetLang}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );
  const reader = response.body.getReader();
  // ... handle SSE events
}

Event Format

translation:

{"sid": 1, "text": "Hello, nice to meet you", "is_final": true}

done:

{"totalUpdated": 10, "characters_billed": 12700, "charged": "6.4", "billed": true}

The three billing fields appear only when the request actually incurred consumption, so the request was billed only when billed is true — always use that field as the criterion. charged is the points consumed by this operation, calculated from the rate, and reflects usage — usage already covered by an unlimited plan is still reported here. Integrators that call this API on behalf of end users and bill them separately can use this value directly instead of deriving it.

error (per-sid sentence translation failure):

When a sentence fails to translate, instead of a translation event, an event: error is sent carrying sid + error_code, interleaved with the translation events. The failure event format matches real-time translation over WebSocket, so the frontend can share one error handler:

event: error
data: {"error_code": "sse_translation_failed", "severity": "error", "message": "SSE translation failed", "context": "sse", "sid": 5, "request_id": "req_abc123xyz789", "timestamp": "2026-04-27T10:30:45.123Z", "details": {"translation_language": "ja-JP", "original_error": "..."}}
FieldTypeDescription
error_codestringError code: sse_translation_failed or llm_content_filtered
severitystringerror for sse_translation_failed; warning for llm_content_filtered
messagestringHuman-readable message
contextstringsse for sse_translation_failed; translation for llm_content_filtered
sidintThe sentence number that failed
request_idstringRequest tracking ID
timestampstringTime the error occurred (ISO 8601)
detailsobjectIncludes debug info such as translation_language and original_error

Failed sentences are saved as translation error records (see the history-playback guide), and the failure markers are visible the next time the history is loaded. For the full specification, see reference/sse/retranslate.md.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Handling
sse_translation_failed500Translation failed (per-sid)The failed sentence is still reported via event: error; the overall flow is not interrupted
llm_content_filtered400The sentence content cannot be translated (per-sid)Retrying will not help; revise the original text and try again. The sentence is excluded from totalUpdated and incurs no consumption
storage_upload_failed500Saving the transcript failedThe entire run is discarded and not billed; retry later. After this code you will not receive done
transcript_revision_conflict409Another write to the same transcript is in progress, the transcript changed while the request was running, or expectedRevision does not match the current revisionThe entire run is discarded and not billed; on a revision mismatch reload the transcript, otherwise simply retry later. After this code you will not receive done
retranslate_segmentation_required422segmented=1 was not sent and the transcript is too long to translate in one request (details carries sentenceCount, maxSentences) (v1.18.0)JSON response before the stream starts, no charge; send segmented=1 to retranslate in segments

Frontend Example

// Use the fetch API to handle SSE
async function retranslate(taskId, targetLang, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/retranslate/${taskId}?targetLang=${targetLang}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    const events = parseSSE(decoder.decode(value));
    for (const event of events) {
      if (event.type === 'translation') {
        console.log(`Sentence ${event.data.sid}: ${event.data.text}`);
      } else if (event.type === 'error') {
        console.warn(`Sentence ${event.data.sid} failed to translate: ${event.data.error_code}`);
      } else if (event.type === 'done') {
        console.log(`Complete, ${event.data.totalUpdated} sentences updated`);
      }
    }
  }
}

GET /api/v1/sse/recordings/{taskId}/entries/{sid}/retranslate (Single-Sentence Retranslation, added in v1.4.0)

Description

Retranslates a single sentence. The most common scenario: after a user edits the original text via PATCH /api/v1/tasks/{id}/entries/{sid}, you call this endpoint to redo all translations for that sentence.

Differences from full-transcript retranslation (/retranslate/{taskId}):

  • Full-transcript retranslation: all sentences are translated into a single target language
  • Single-sentence retranslation: only one sentence is translated, and every language it has been translated into or failed to translate into can be retried at once; supports optimistic locking

Authentication

Query: api_key (the browser EventSource does not support headers)

Request Parameters

ParameterLocationTypeRequiredDescription
taskIdpathstringYesRecording ID (UUID)
sidpathnumberYesSentence ID (1-based)
targetLangquerystringNoTarget language code. When omitted, every language that sentence has either been translated into or failed to translate into is retried, that is the union of the translations and the translation error records
expectedRevisionquerynumberNoOptimistic lock: the current transcript revision; a mismatch returns transcript_revision_conflict
api_keyquerystringYesAPI key

Event Format

Event sequence: connected → progress / translated / error ×N → done

// progress (when translation begins for each language)
{ "sid": 5, "lang": "en-US", "status": "translating" }

// translated (when each language completes successfully)
{ "sid": 5, "lang": "en-US", "text": "Hello world", "tokens_used": 25 }

// done (all complete; successfully translated languages are listed in languages_translated)
{
  "sid": 5,
  "revision": 6,
  "original_text_edited_at": "2026-05-06T10:30:00.000000Z",
  "languages_translated": ["en-US"],
  "languages_failed": ["ja-JP"]
}

Error Responses

Error CodeHTTPDescription
recording_not_found404Recording does not exist or does not belong to the user
recording_not_completed422The recording has not finished processing
entry_not_found404The specified sentence was not found
entry_text_empty422The original text of that sentence is empty (whitespace-only counts as empty)
sse_translation_failed500A target language failed to translate (per-lang)
llm_content_filtered400A target language's content cannot be translated (per-lang); retrying will not help
transcript_revision_conflict409Revision mismatch, or another write to the same transcript is in progress
storage_upload_failed500Failed to save the transcript

For the full event format and a workflow example combining optimistic locking with PATCH, see reference/sse/retranslate.md.


init_sentence Edit Marker Fields (added in v1.4.0)

For sentences edited by a user, historyTranscribe adds two fields to the init_sentence event (only present after editing):

{
  "sid": 7,
  "origin": "Corrected text",
  "original_text_raw": "Original STT output",
  "original_text_edited_at": "2026-05-06T10:30:00.000000Z",
  "translations": { "en-US": "Corrected text" }
}

Frontend detection: determine this by the presence of the field ('original_text_raw' in data); do not compare origin === original_text_raw — a user may edit and then change it back to the same string, in which case the text is equal but the "edited" marker should still be shown. See reference/sse/history.md.


GET /api/v1/sse/retranslate/summary/{taskId} (Retranslate Summary)

Description

Retranslates the summary of the specified task into the target language. Translation results are delivered segment by segment via an SSE stream.

Use Cases

  • Switching the summary display language
  • Obtaining the summary in a different language

Authentication

Header: X-API-Key: YOUR_API_KEY

Request Parameters

ParameterTypeRequiredDescription
taskIdstringYesRecording ID
targetLangstringYesTarget language code

Request Example

// Use the fetch API (because EventSource does not support headers)
async function retranslateSummarySSE(taskId, targetLang, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/retranslate/summary/${taskId}?targetLang=${targetLang}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );
  const reader = response.body.getReader();
  // ... handle SSE events
}

Event Format

summary_translation:

{"text": "Accumulated translation result...", "is_final": false}

done:

{"totalUpdated": 1}

The retranslated summary is not saved: the saved summary and its language are unchanged. To switch the summary to another language and keep it, use the save endpoint of Regenerate Summary (POST, billed).

When the translation is incomplete, done also carries truncated: true (added in v1.17.0). Processing time and timeouts follow the same rules as Summary Translation.

This endpoint is not billed. Its done event does not include the characters_billed / charged / billed fields.

Error Responses

Error CodeHTTP StatusDescriptionRecommended Handling
sse_summary_not_found404Summary not foundThis recording has no summary
sse_summary_translation_failed500Summary translation failed. details.original_error of Translation timed out means the request timed outRetry later
llm_content_filtered400The summary content cannot be translatedRetrying will not help; revise the summary and try again

Regenerate Summary (GET Preview / POST Save)

Split into two endpoints + mode-aware. For the full schema, see reference/sse/regenerate-summary.md; this is a quick summary.

MethodEndpointPersists ResultSaves TranscriptBilledPurpose
GET/api/v1/sse/regenerate/summary/{taskId}NoNoYesPreview (dry run)
POST/api/v1/sse/regenerate/summary/{taskId}YesYes (and increments revision)YesSave (persist officially)

Known limitation: GET is also billed — the LLM actually consumes tokens, so the GET endpoint cannot be used for free.

Shared Parameters (GET via query string, POST via JSON body)

ParameterTypeRequiredDescription
taskId (path)stringYesRecording UUID
modestringYesSummary mode enum: builtin / custom
templatestringRequired for builtin / forbidden for customBuilt-in template slug
promptstringRequired for custom / forbidden for builtinThe customer's full prompt (replaces the built-in layered prompt, ≤3000 characters)
promptSlugstringRequired for custom / forbidden for builtinThe customer's own identifier (≤64 Unicode characters, no control characters)
languagestringNoOutput language (defaults to the first transcription language)
plainTextbooleanNoWhether to request plain-text output (default false)

Mutual exclusivity rule: a violation is a parameter validation failure (HTTP 200 + an error event whose data carries only message, with no error_code).

Request Example

# Preview builtin (does not persist the result)
curl -N "https://vas-poc.vurbo.ai/api/v1/sse/regenerate/summary/550e8400-...?mode=builtin&template=meeting&language=zh-TW&plainText=true" \
  -H "X-API-Key: YOUR_API_KEY"

# Save custom
curl -N -X POST "https://vas-poc.vurbo.ai/api/v1/sse/regenerate/summary/550e8400-..." \
  -H "X-API-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"mode":"custom","prompt":"Please emphasize KPIs","promptSlug":"acme-v2","plainText":true}'

Event Sequence

1. connected              → Connection confirmation (includes mode=builtin|custom, endpoint=preview|persist)
2. summary_regeneration   → Stream summary segments (accumulating; is_final=true marks the last one)
3. done                   → Complete, includes final_content / mode / template(effective) / prompt_snapshot (only for custom)
   or
3. error                  → Generation failed (sse_summary_regeneration_failed; can occur on both endpoints;
                            nothing is saved, no done is sent, and it is not billed)
   or
3. error                  → Save failed, or another write is in progress (save endpoint only; the request
                            stops, no done is sent, and it is not billed)

done event

{
  "task_id": "550e8400-...",
  "tokens_used": 123,
  "final_content": "This meeting...",
  "mode": "custom",
  "template": "acme-v2",
  "plain_text": true,
  "persisted": true,
  "summary_language": "zh-TW",
  "characters_billed": 12700,
  "charged": "1.3",
  "billed": true,
  "prompt_snapshot": "Please emphasize KPIs"
}
  • mode: summary mode (builtin / custom)
  • template: effective slug — builtin → built-in template slug; custom → customer slug
  • persisted: whether this summary has been officially saved (false for GET, true for POST)
  • summary_language: the language actually used for this summary. The language value if one was sent; otherwise the first transcription language. Present for both preview (GET) and save (POST), and always has a value
  • characters_billed / charged / billed: consumption for this request. Present only when consumption actually occurred (absent when generation fails); always use billed to determine whether the request was billed. Both preview (GET) and save (POST) are billed. charged reflects usage — usage already covered by an unlimited plan is still reported here
  • prompt_snapshot: only present in custom mode; the prompt content the customer passed in verbatim (a mandatory snapshot, the sole basis for reconstruction)
  • truncated: present only when the summary is incomplete (the value is always true): the summary reached the output length limit, or generation reached the processing time limit (only the part completed so far is returned). Billed as usual, and the save endpoint also saves this summary

Error Codes

Error CodeHTTPDescription
recording_not_found404Recording not found
sse_template_not_found404Summary template not found
sse_transcript_not_found404Transcript not found
summary_text_empty400The transcript has no content
summary_text_too_long400The transcript exceeds the 200,000-character limit
sse_summary_regeneration_failed500Regeneration failed (the response does not include internal error details; content filtering, or a stream that stalls or does not end normally, also falls under this code; nothing is saved or billed)

Parameter validation failures carry no error code. An invalid mode, a field combination that does not match the mode, a prompt or promptSlug that is too long or contains control characters — all of these are rejected during parameter validation. The response is an error event whose data carries only message, with no error_code. Present message to the user; do not try to match on an error code.

(The set_summary action for realtime recording takes a different path, and it does have dedicated error codes — see the WebSocket API documentation.)

Frontend Example

async function regenerateSummary(taskId, body, apiKey, { persist = false } = {}) {
  const url = `https://vas-poc.vurbo.ai/api/v1/sse/regenerate/summary/${taskId}`;
  const init = persist
    ? { method: 'POST', headers: { 'X-API-Key': apiKey, 'Content-Type': 'application/json' }, body: JSON.stringify(body) }
    : { method: 'GET', headers: { 'X-API-Key': apiKey } };
  if (!persist) {
    const params = new URLSearchParams(body);
    return fetch(`${url}?${params}`, init);
  }
  return fetch(url, init);
}

POST /api/v1/sse/summary (Ad-hoc Summary, added in v1.9.1)

Generates a summary from arbitrary text supplied in the request (not tied to a recording). For the full schema, see reference/sse/adhoc-summary.md; this is a quick summary.

MethodEndpointPersists ResultBilledPurpose
POST/api/v1/sse/summaryNoYesGenerate a summary from the content supplied in the request (content the server does not have, such as merged full transcripts or edited transcripts); the result is not stored
  • Required: content (≤200,000 characters), idempotency_key (≤64 characters, A-Z a-z 0-9 . _ -), mode (builtin / custom; the field mutual-exclusion rules are the same as Regenerate Summary)
  • Billing: 0.1 credits per 1,000 content characters (same rate as summaries); billed only on successful generation
  • Duplicate-request guarantee (idempotency_key is only valid within a single API key): the entire request (content plus every summary parameter) determines whether a call is a retry. Retrying with the same identifier + an identical request is not billed again (but regenerates — the previous result is not replayed); the same identifier + any differing field returns 409 summary_idempotency_key_conflict; failures do not claim the identifier
  • Error contract differs from other SSE endpoints: before the stream starts, real HTTP status codes are returned (401/403 / 422 / 404 / 400 / 402 / 409), and authentication only accepts the header — ?api_key= is not accepted; SSE error events are only used after the stream starts
  • Event sequence: connected → summary_regeneration ×N → done (the done event has no task_id / persisted, and adds the characters_billed / charged / idempotency_key / billed fields; summary_language is the same as in Regenerate Summary and is zh-TW when language is not sent). When the summary is incomplete, done also carries truncated: true (billed as usual); when generation fails, error (sse_summary_regeneration_failed) is sent instead and nothing is billed
  • Billing fields: characters_billed and charged are always present; billed is false for free retries and empty generation results — reconcile against that field. charged reflects usage — usage already covered by an unlimited plan is still reported here

POST /api/v1/sse/summary/translate (Summary Translation, added in v1.17.0)

Translates summary text supplied in the request into the specified language, without tying it to a recording. See reference/sse/summary-translate.md for the full specification; this is a quick overview.

MethodEndpointResult savedBilledPurpose
POST/api/v1/sse/summary/translateNoYesTranslate the content supplied in the request, such as a merged or edited summary the server does not have. The result is not saved
  • Required:
    • content: 30,000 characters or fewer
    • target_language
    • idempotency_key: 64 characters or fewer, using only A-Z a-z 0-9 . _ -
  • Optional: source_language. Detected automatically when omitted. Returns 422 if it matches target_language.
  • Billing: 0.1 credits per 200 characters of content; a partial unit counts as a full one. This is the same rate as full-transcript retranslation. Billed only on success.
  • Duplicate requests: same rules as Ad-hoc Summary.
    • Retrying the identical request with the same identifier is not billed again.
    • Reusing the identifier with any field changed returns 409 summary_idempotency_key_conflict.
  • Error contract: errors before the stream starts return a real HTTP status code (401/403, 422, 402, 409, 429). Only errors after the stream starts use an SSE error event.
  • Event sequence: connected → summary_translation ×N (cumulative full text) → done.
    • For long content, the stream often delivers a large block at once, and there may be a pause of a few seconds between events.
  • done fields: tokens_used, source_language, target_language, characters_billed, charged, idempotency_key, billed.
    • An incomplete translation also carries truncated: true: it either hit the processing-time or length limit and was cut off (billed as usual), or was judged unfinished (not billed). Check billed to see whether credits were actually deducted.
  • Processing time: a request has a limit of about 230 seconds. If translation pauses for more than 60 seconds, an error is sent with error code sse_summary_translation_failed and details.original_error set to Translation timed out.

GET /api/v1/sse/tts/{taskId} (TTS Audio Stream)

Description

Converts the translated content of a historical recording into TTS audio, delivered sentence by sentence via an SSE stream. The frontend can control how many sentences are returned per request.

Use Cases

  • Audio playback of translations from historical recordings
  • Karaoke effect (combined with word boundaries)
  • Voice readout of translated content

Authentication

Header: X-API-Key: YOUR_API_KEY

Request Parameters

ParameterTypeRequiredDescription
taskIdstringYesRecording ID (path parameter)
languagestringYesTTS output language (e.g., en-US)
voicestringNoSpecify a voice name (e.g., en-US-JennyNeural)
sidintNoStarting sentence ID (default 1, starting from the first sentence)
lengthintNoNumber of sentences to return (default 1, maximum 20)

Note: The maximum value of length is 20 by default and is configured on the server side. It is automatically truncated when it exceeds the maximum.

Request Example (Single Sentence Playback)

// Use the fetch API (because EventSource does not support headers)
async function playTTSSingle(taskId, language, sid, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/tts/${taskId}?language=${language}&sid=${sid}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );
  const reader = response.body.getReader();
  // ... handle SSE events
}

Request Example (Multiple Sentence Playback)

// Play sentences 5, 6, and 7 (3 sentences total)
async function playTTSMultiple(taskId, language, sid, length, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/tts/${taskId}?language=${language}&sid=${sid}&length=${length}`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );
  const reader = response.body.getReader();
  // ... handle SSE events
}

Event Sequence

1. connected    → Connection confirmation
2. tts_audio    → Send TTS audio sentence by sentence (repeated N times, N = length)
3. tts_done     → Playback complete

Event Format

connected:

{
  "task_id": "550e8400-e29b-41d4-a716-446655440000",
  "language": "en-US",
  "voice": "en-US-JennyNeural",
  "start_sid": 5,
  "length": 3
}

tts_audio:

{
  "sid": 5,
  "transcript": "Hello, nice to meet you",
  "text": "Hello, nice to meet you",
  "audio": "Base64EncodedMP3...",
  "duration_ms": 2500,
  "boundaries": [
    {"offset_ms": 0, "duration_ms": 350, "text_offset": 0, "word_length": 5, "text": "Hello"},
    {"offset_ms": 350, "duration_ms": 100, "text_offset": 5, "word_length": 1, "text": ","},
    {"offset_ms": 500, "duration_ms": 250, "text_offset": 7, "word_length": 4, "text": "nice"},
    {"offset_ms": 750, "duration_ms": 200, "text_offset": 12, "word_length": 2, "text": "to"},
    {"offset_ms": 950, "duration_ms": 350, "text_offset": 15, "word_length": 4, "text": "meet"},
    {"offset_ms": 1300, "duration_ms": 300, "text_offset": 20, "word_length": 3, "text": "you"}
  ],
  "characters_used": 23
}
FieldTypeDescription
sidintSentence ID
transcriptstringOriginal transcript (STT recognition result)
textstringTranslated text (the TTS synthesis source)
audiostringBase64-encoded MP3 audio
duration_msintAudio duration (milliseconds)
boundariesarrayWord boundary array
characters_usedintCharacters consumed synthesizing this sentence; 0 on a cache hit (tts_done.total_characters_used is the sum of this field)

Word Boundary Field Descriptions

FieldTypeDescription
offset_msintThe word's start time in the audio (milliseconds)
duration_msintThe word's duration (milliseconds)
text_offsetintPosition in the original text string (character index)
word_lengthintWord length (number of characters)
textstringWord content

tts_done:

{
  "sentences_sent": 3,
  "total_duration_ms": 7500,
  "total_characters_used": 142
}
FieldTypeDescription
sentences_sentintThe number of sentences actually sent
total_duration_msintThe total audio duration of all sentences (milliseconds)
total_characters_usedintThe total number of characters synthesized in this TTS request (used for quota calculation)

Error Responses

Error CodeHTTP StatusDescriptionRecommended Handling
recording_not_found404Recording not foundConfirm the taskId is correct
tts_synthesis_failed500TTS synthesis failedRetry later

Frontend Example

// Use the fetch API to handle TTS SSE
async function playTTS(taskId, language, apiKey, startSid = 1, length = 1) {
  const url = new URL(`https://vas-poc.vurbo.ai/api/v1/sse/tts/${taskId}`);
  url.searchParams.set('language', language);
  url.searchParams.set('sid', startSid);
  url.searchParams.set('length', length);

  const response = await fetch(url, {
    headers: {
      'X-API-Key': apiKey
    }
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    const events = parseSSE(decoder.decode(value));
    for (const event of events) {
      if (event.type === 'connected') {
        console.log(`TTS connection successful, voice: ${event.data.voice}`);
      } else if (event.type === 'tts_audio') {
        console.log(`Sentence ${event.data.sid}: ${event.data.text}`);

        // Play the audio
        const audioBlob = base64ToBlob(event.data.audio, 'audio/mp3');
        const audioUrl = URL.createObjectURL(audioBlob);
        const audio = new Audio(audioUrl);

        // Set up the karaoke effect
        setupKaraoke(audio, event.data.boundaries, event.data.text);

        audio.play();
      } else if (event.type === 'tts_done') {
        console.log(`Playback complete, ${event.data.sentences_sent} sentences total`);
      }
    }
  }
}

// Base64 to Blob
function base64ToBlob(base64, mimeType) {
  const byteCharacters = atob(base64);
  const byteNumbers = new Array(byteCharacters.length);
  for (let i = 0; i < byteCharacters.length; i++) {
    byteNumbers[i] = byteCharacters.charCodeAt(i);
  }
  const byteArray = new Uint8Array(byteNumbers);
  return new Blob([byteArray], { type: mimeType });
}

// Karaoke effect
function setupKaraoke(audio, boundaries, text) {
  const updateHighlight = () => {
    const currentTimeMs = audio.currentTime * 1000;
    const currentWord = boundaries.find((b, i) => {
      const nextOffset = boundaries[i + 1]?.offset_ms ?? Infinity;
      return currentTimeMs >= b.offset_ms && currentTimeMs < nextOffset;
    });

    if (currentWord) {
      // Highlight the current word
      highlightWord(text, currentWord.text_offset, currentWord.word_length);
    }
  };

  const interval = setInterval(updateHighlight, 50);
  audio.addEventListener('ended', () => clearInterval(interval));
}

GET /api/v1/sse/imports/{importId}/progress (Import Progress Stream)

Description

Tracks the processing progress of an audio file import in real time. After connecting, progress updates are continuously pushed via an SSE stream until the import completes, fails, or the connection times out.

Use Cases

  • Showing a real-time processing progress bar after uploading an audio file
  • Tracking the progress of each stage: audio conversion, transcription, translation, summary, etc.

Authentication

Header: X-API-Key: YOUR_API_KEY

Request Parameters

ParameterTypeRequiredDescription
importIdstringYesImport task ID (UUID, path parameter)

Request Example

curl -N "https://vas-poc.vurbo.ai/api/v1/sse/imports/550e8400-e29b-41d4-a716-446655440000/progress" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"

Event Sequence

Scenario 1: Import still in progress
1. connected       → Connection confirmation
2. progress        → Send the current progress
3. progress ×N     → Continuously pushed when progress changes
   heartbeat ×N    → Sent every 15 seconds when there is no progress change
4. completed       → Import succeeded, connection ends
   or failed       → Import failed, connection ends
   or timeout      → Exceeded 15 minutes, connection ends

Scenario 2: Import already complete (terminal state)
1. connected       → Connection confirmation
2. progress        → Send the final progress
3. completed       → Send the completed event directly and end
   or failed       → Send the failed event directly and end

Event Format

connected:

{"message": "Import progress service connected (importId: xxx)"}

progress:

{
  "import_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "processing",
  "stage": "transcribing",
  "progress": 45,
  "message": "Transcribing..."
}
FieldTypeDescription
import_idstringImport task ID (UUID)
statusstringImport status: pending / processing / completed / failed
stagestring / nullThe current processing stage
progressintegerProgress percentage (0-100)
messagestringHuman-readable progress message

Stage values and their corresponding progress ranges:

ValueDescriptionProgress Range
convertingAudio format conversion0% - 10%
transcribingSpeech-to-text10% - 60%
translatingText translation60% - 85%
summarizingGenerating the summary85% - 100%
completedImport complete100%
nullNot started yet—

completed:

{
  "import_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "completed",
  "task_id": "abc123-e29b-41d4-a716-446655440000",
  "message": "Processing complete"
}
FieldTypeDescription
import_idstringImport task ID
statusstringFixed as completed
task_idstringThe generated task ID, which can be used for subsequent queries
messagestringFixed as Processing complete

failed:

{
  "import_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "failed",
  "error_code": "import_invalid_format",
  "error_message": "Unsupported audio format"
}
FieldTypeDescription
import_idstringImport task ID
statusstringFixed as failed
error_codestringError code
error_messagestringHuman-readable error message (a general description without internal details; provide the import_id for troubleshooting)

heartbeat:

Sent every 15 seconds when there is no progress change, used to keep the connection alive.

{"timestamp": 1708761600}

timeout:

Sent when the import has not completed after 15 minutes; the connection ends automatically.

{"message": "Connection timeout"}

Error Responses

Error CodeHTTP StatusDescriptionRecommended Handling
import_not_found404The specified import task was not foundConfirm the importId is correct

Frontend Example

async function trackImportProgress(importId, apiKey) {
  const response = await fetch(
    `https://vas-poc.vurbo.ai/api/v1/sse/imports/${importId}/progress`,
    {
      headers: {
        'X-API-Key': apiKey
      }
    }
  );

  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    buffer += decoder.decode(value, { stream: true });
    const events = buffer.split('\n\n');
    buffer = events.pop();

    for (const eventStr of events) {
      if (!eventStr.trim()) continue;

      const lines = eventStr.split('\n');
      let eventType = '';
      let eventData = '';

      for (const line of lines) {
        if (line.startsWith('event: ')) eventType = line.slice(7);
        else if (line.startsWith('data: ')) eventData = line.slice(6);
      }

      if (!eventType || !eventData) continue;
      const data = JSON.parse(eventData);

      switch (eventType) {
        case 'connected':
          console.log('Connected:', data.message);
          break;
        case 'progress':
          console.log(`[${data.stage}] ${data.progress}% - ${data.message}`);
          updateProgressBar(data.progress, data.stage, data.message);
          break;
        case 'completed':
          console.log('Import complete! Recording ID:', data.task_id);
          navigateToRecording(data.task_id);
          break;
        case 'failed':
          console.error('Import failed:', data.error_code, data.error_message);
          showError(data.error_message);
          break;
        case 'timeout':
          console.warn('Connection timeout:', data.message);
          break;
      }
    }
  }
}

Version: V1.24.1 Last Updated: 2026-09-28

Copyright © 2026