History SSE
Connection Info
| Item | Value |
|---|---|
| Base path | https://vas-poc.vurbo.ai/api/v1/sse |
| Protocol | HTTP + Server-Sent Events (SSE) |
| Data format | text/event-stream |
| Auth method | Header X-API-Key: {KEY}, or query ?api_key={KEY} (either works; query takes precedence) |
Note: The browser's native EventSource API does not support custom headers. Use the fetch API with ReadableStream, or use an SSE client library that supports headers.
Endpoint Overview
| Method | Endpoint | Description |
|---|---|---|
| GET | /api/v1/sse/history/transcribe/{taskId} | Retrieve historical conversation records |
GET /api/v1/sse/history/transcribe/{taskId}
Description
Loads the complete conversation record for a specified task, including all sentences and the summary. The data is sent one item at a time over an SSE stream.
Difference from the Transcript Download API (
GET /api/v1/tasks/{taskId}/transcript/export):
- This endpoint: for progressive loading; pushes raw structured data (JSON fragments) sentence by sentence as an event stream, so the front end can render the UI progressively.
- Transcript download: for offline download; returns the complete file (TXT / SRT / SBV / VTT / CSV) in one response, ready to open in subtitle software or a spreadsheet.
Use Cases
- View the recording details page
- Load historical transcripts
Authentication
Header: X-API-Key (see Authentication)
Request Parameters
| Parameter | Location | Type | Required | Description |
|---|---|---|---|---|
taskId | path | string | Yes | Recording ID (UUID) |
Request Example
curl -N "https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/550e8400-e29b-41d4-a716-446655440000" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"
// Use the fetch API (because EventSource does not support headers)
async function connectSSE(taskId, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/${taskId}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
// ... handle SSE events
}
Event Sequence
1. connected → connection confirmation
2. init_metadata → send task metadata
3. init_sentence → send sentences one by one (repeats N times)
4. init_summary → send summary
5. init_done → initialization complete
Event Formats
connected
{
"message": "History service connected (taskId: xxx)"
}
init_metadata
{
"task_id": "550e8400-e29b-41d4-a716-446655440000",
"title": "Meeting Notes",
"created_at": "2026-02-23T10:00:00Z",
"type": "transcribe",
"has_speaker_diarization": true,
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"summary_template": "general",
"summary_language": "zh-TW",
"speaker_aliases": {"speaker_1": "Manager Wang"}
}
| Field | Type | Description |
|---|---|---|
task_id | string | Task ID (UUID) |
title | string | Task title |
created_at | string | Creation time (ISO 8601) |
type | string | Recording type |
has_speaker_diarization | boolean | Whether speaker diarization is enabled |
transcription_languages | array | Array of transcription languages (BCP 47, e.g. ["zh-TW"]), up to 10; [] when there is no data |
translation_languages | array | Array of translation languages (BCP 47, e.g. ["en-US", "ja-JP"]), up to 12; [] when there is no data |
summary_template | string|null | Summary template slug (e.g. general, meeting); null if not specified |
summary_language | string|null | Summary language (BCP 47, e.g. zh-TW, en-US). If no summary language was specified for the recording, this is usually the first transcription language; in two-way mode with active_language specified, it is that language. It may be null for tasks without a summary. To determine the language of the summary itself, use the summary_language of init_summary |
speaker_aliases | object | Mapping of "original speaker ID → display name"; {} (an empty object, not an array) when there are no aliases. The front end uses this for duplicate-name precheck before a rename (added in v1.3.12) |
init_sentence
{
"sid": 1,
"origin": "Hello",
"translations": {
"en-US": "Hello"
},
"start_time": "00:05",
"speaker_id": "speaker_1",
"speaker_label": "Manager Wang"
}
If a sentence has a translation failure, it carries an additional translation_errors field (present only when there is a failure). This lets the front end distinguish "the language was never scheduled for translation" (translations missing the key) from "it was translated but failed" (translation_errors has the key). The same language may have both an older translation and a failure record (a failed retranslation keeps the previous translation), so read both fields together:
{
"sid": 5,
"origin": "Sentence with sensitive words",
"translations": {
"en-US": "Sensitive sentence"
},
"translation_errors": {
"ja-JP": "llm_content_filtered"
},
"start_time": "00:25",
"speaker_id": "speaker_1",
"speaker_label": "Manager Wang"
}
If the user has edited the original text of a sentence (via PATCH /api/v1/tasks/{id}/entries/{sid}), it carries additional original_text_raw and original_text_edited_at fields (present only after editing):
{
"sid": 7,
"origin": "Corrected text",
"original_text_raw": "Original STT output",
"original_text_edited_at": "2026-05-06T10:30:00.000000Z",
"translations": { "en-US": "Corrected text" },
"start_time": "00:35",
"speaker_id": "speaker_1",
"speaker_label": "Manager Wang"
}
For multi-channel recordings (recognition_mode: "multi_channel"), each sentence carries an additional channel_id field (present only for multi-channel recordings; single-channel recordings do not carry this field):
{
"sid": 9,
"origin": "Good morning, everyone",
"translations": {
"en-US": "Good morning, everyone"
},
"start_time": "00:40",
"speaker_id": "channel_1",
"speaker_label": "Manager Wang",
"channel_id": 1
}
| Field | Type | Description |
|---|---|---|
sid | number | Sentence ID |
origin | string | Original text content (the user-corrected version if it has been edited) |
translations | object|null | Map of translated text ({"language_code": "translated text"}); null when there is no translation |
translation_errors | object | Optional. Map of translation failure error codes ({"language_code": "error_code"}); this field is omitted when there are no failures |
original_text_raw | string | Optional. The raw STT output text. Present only when the user has edited the sentence. The front end can use it to display an "edited" marker and offer a "restore original text" function |
original_text_edited_at | string | Optional. The most recent edit time of the original text (ISO 8601). Appears together with original_text_raw |
start_time | string | Start time (mm:ss) |
speaker_id | string|null | The speaker ID (e.g. Guest-1; the same person is not guaranteed to keep the same ID across a connection recovery — see Speaker Management). Provided as the source for target_speaker_id in PATCH /speakers/reassign (v1.5.3 reversal: previously the display name) |
speaker_label | string|null | The display label (the human-readable name after applying speaker_aliases, e.g. Manager Wang). Equal to speaker_id when there is no alias (added in v1.5.3 to replace the former display semantics of speaker_id) |
channel_id | number | Optional. The source channel number of this sentence (the value of channels[].channel_id at start). Present only for multi-channel recordings; single-channel recordings do not carry this field (added in V1.10.0) |
Front-end detection: Determine whether a sentence has been edited by the presence of the field (
'original_text_raw' in dataordata.original_text_raw !== undefined). Do not compareorigin === original_text_raw— the user may have edited the text and then changed it back to the same string; in that case the text is equal but the "edited" marker should still be shown.
v1.5.3 naming reversal:
speaker_idis reversed from the display name to the original ID; a newspeaker_labelfield holds the display label. Speaker edits (reassign/merge) always usespeaker_idas the locating key. See the V1.5.3 changelog.
Multi-channel recordings (V1.10.0):
speaker_iduses thechannel_{N}format (N = channel number); speaker identity is determined by the physical channel rather than inferred by AI.speaker_labelis thespeaker_nameconfigured for that channel (or the display name after arename_speakerchange); it is equal tospeaker_idwhen there is no alias. The informational overlap betweenchannel_idandspeaker_idis deliberate —channel_idis an immutable record of the source channel, so each can be traced separately for auditing. Also note that for multi-channel recordings, sentences are sent sorted by time andsidis not guaranteed to be increasing; clients must not usesidvalues to infer ordering.
init_summary
In addition to the summary text text, this includes mode-aware metadata (mode / template / plain_text / prompt_snapshot), which lets the client trace the mode, effective slug, and customer prompt content (custom mode) that correspond to that summary.
v1.5.5 adds fallback_level / dropped_segments: these appear only when the summary actually went through the LLM service content-filter automatic fallback (neutral mode or segment-omission mode), for auditing and UI hints during history playback.
Example (standard mode succeeds directly, no fallback):
{
"text": "Summary content...",
"mode": "custom",
"template": "acme-meeting-v2",
"plain_text": true,
"summary_language": "zh-TW",
"prompt_snapshot": "Please emphasize KPIs"
}
Example (segment-omission mode triggered, generated after 2 transcript segments were omitted):
{
"text": "Summary content (2 segments omitted)...",
"mode": "custom",
"template": "acme-meeting-v2",
"plain_text": true,
"summary_language": "zh-TW",
"prompt_snapshot": "Please emphasize KPIs",
"fallback_level": 3,
"dropped_segments": [3, 7]
}
| Field | Type | Description |
|---|---|---|
text | string | Summary text |
mode | string | null | "builtin" / "custom" / null (null when no summary was generated) |
template | string | null | effective slug — builtin → built-in template slug; custom → customer slug |
plain_text | boolean | Whether the output is plain text |
summary_language | string | null | The language of this summary (BCP 47, e.g. zh-TW). Always present. If no summary language was specified for the recording, this is usually the first transcription language; in two-way mode with active_language specified, it is that language. null when there is no summary (text is an empty string) (added in v1.16.5) |
prompt_snapshot | string | null | Has a value only in custom mode; the prompt content passed in verbatim by the customer (the basis for reconstruction) |
fallback_level | int (omit) | Present only when a fallback was triggered (2 or 3). 2 = neutral mode (regenerated with a neutral instruction); 3 = segment-omission mode (regenerated after the offending segments were omitted). Omitted when standard mode succeeds directly |
dropped_segments | int[] (omit) | Present only when fallback_level=3; the indices of the trimmed transcript segments (an integer array in original order) |
summary_languageand the field of the same name ininit_metadata: to determine the language of the summary itself, use thesummary_languageof this event. When there is no summary,init_metadata.summary_languagemay still carry the language setting from the recording while this field isnull; the two differing is expected.
fallback_level/dropped_segmentsandprompt_snapshotare complementary: the former records the actual execution path (whether a fallback was taken), and the latter records the customer intent (the original prompt content). Even if a fallback was triggered and the customer prompt was not actually used,prompt_snapshotstill preserves the original text as an audit record. See V1.5.5 changelog – LLM service content-filter automatic fallback.
init_done
{
"totalSentences": 10
}
| Field | Type | Description |
|---|---|---|
totalSentences | number | Total number of sentences |
Edge Case: No Speech Content (V1.3.7)
If the task is silent throughout, the volume is too low, there is too much noise, or the recognition language does not match the actual audio—so that the speech recognition engine recognizes no sentences—this endpoint still completes with the normal event sequence (not an sse_transcript_not_found error):
init_metadatais sent normallyinit_sentenceis sent 0 times (no sentences)- The
textofinit_summaryis an empty string"", and itssummary_languageisnull - The
totalSentencesofinit_doneis0
This behavior applies to both sources: real-time recording (WebSocket recording ends) and file import (offline processing completes), and is aligned with the "zero recognition results" legalization behavior of the V1.3.5 import flow. The client should use totalSentences === 0 to decide whether to show a "no speech content" empty state, rather than treating it as an error branch. See File Import Guide – Behavior When Audio Cannot Be Recognized.
Specific Error Codes
| Error Code | HTTP Status | Description | Recommended Action |
|---|---|---|---|
recording_not_found | 404 | Recording not found | Verify that taskId is correct |
sse_transcript_not_found | 404 | Transcript not found | The transcript for the specified taskId does not exist or could not be read. A real-time recording that received no audio at all (task_complete carried noAudio: true) has no transcript and returns this error; a recording that captured audio but no speech does not |
Front-End Example
async function loadHistory(taskId, apiKey) {
const response = await fetch(
`https://vas-poc.vurbo.ai/api/v1/sse/history/transcribe/${taskId}`,
{
headers: {
'X-API-Key': apiKey
}
}
);
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const text = decoder.decode(value);
// Parse the SSE format: event: xxx\ndata: {...}\n\n
const events = parseSSE(text);
for (const event of events) {
if (event.type === 'init_metadata') {
console.log('Task info:', event.data.title);
} else if (event.type === 'init_sentence') {
console.log(`[${event.data.start_time}] ${event.data.origin}`);
if (event.data.translations) {
console.log(`Translations:`, event.data.translations);
}
} else if (event.type === 'init_summary') {
console.log('Summary:', event.data.text);
} else if (event.type === 'init_done') {
console.log(`Load complete, ${event.data.totalSentences} sentences total`);
}
}
}
}
Version: V1.24.1 Last Updated: 2026-10-01