Floating Subtitle SSE
Overview
The Floating Subtitle SSE lets you subscribe, over an independent, read-only connection, to the live transcript of an in-progress recording (source-language original text plus target-language translations). It is designed for desktop floating-subtitle windows, second-screen captions, and similar use cases. This feed is separate from the recording's own WebSocket connection, so it can be opened independently on a different device or window. "In-progress recordings" include live broadcasts — a broadcast host can also use this feed to subscribe to their own live transcript during a broadcast.
Note: The base path for Floating Subtitle SSE is
https://vas-poc.vurbo.ai(the real-time service, same origin as the WebSocket), which differs from the general SSE API basehttps://vas-poc.vurbo.ai/api/v1/sse.
Connection Information
| Item | Value |
|---|---|
| Base path | https://vas-poc.vurbo.ai |
| Protocol | HTTP + Server-Sent Events (SSE) |
| Data format | text/event-stream |
| Authentication | feed_token (bound to the recording, no API Key required) |
Endpoint Overview
| Method | Endpoint | Description |
|---|---|---|
| POST | /api/v1/auth/tasks/{taskId}/subtitle-feed-token | Exchange for a floating-subtitle feed_token (see Subtitle Feed Token API) |
| GET | /tasks/{task_id}/subtitle | Real-time transcript SSE stream |
Connection Flow
- The recording owner calls
POST /api/v1/auth/tasks/{taskId}/subtitle-feed-tokenwith an API Key to exchange for a short-livedfeed_token. - The floating-subtitle window connects to
GET /tasks/{task_id}/subtitlewith thatfeed_tokento receive the SSE stream. - The
feed_tokencan be reused for reconnection while valid (each successful validation extends its validity); for long recordings, exchange for a new token before expiry.
On-site audience members can also connect via the owner's share link: exchange the share secret for an audience Token (see Subtitle Feed Token API), then connect to this SSE with that Token.
GET /tasks/{task_id}/subtitle
Description
Subscribe read-only to the live transcript of an in-progress recording, receiving an SSE stream of original text (both interim and finalized) and translation results. Supports cross-device and cross-window connections, with automatic reconnection and replay on reconnect.
Use Cases
- Persistent desktop floating-subtitle window
- Second-screen / projected caption display
- Bilingual (original + translation) real-time subtitles
Authentication
feed_token authentication (no API Key required): verified via the feed_token query parameter. The token is bound to the recording and uses a sliding expiry (each successful validation extends it).
Request Parameters
| Parameter | Location | Type | Required | Description |
|---|---|---|---|---|
task_id | path | string | Yes | Recording ID |
feed_token | query | string | Yes | Floating-subtitle access Token (obtained from the token endpoint) |
lang | query | string | No | Filter target translation languages; comma-separated (e.g., en-US,ja-JP). Omit or * for all |
Request Examples
// Receive all languages
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/tasks/3f9a.../subtitle?feed_token=xxx'
);
// Receive only English translation
const eventSource = new EventSource(
'https://vas-poc.vurbo.ai/tasks/3f9a.../subtitle?feed_token=xxx&lang=en-US'
);
Connection Timing Boundaries
| Situation | HTTP Status | Recommended Handling |
|---|---|---|
| Recording has not started (or timed out and disappeared) | 425 Too Early | Retry later |
| Recording has ended | 410 Gone | Inform the user and close the window |
| feed_token invalid / expired | 401 Unauthorized | Obtain a new feed_token |
| Connection or audience limit reached | 429 Too Many Requests | Retry later or close extra windows |
Event Types
| Event | Description | Notes |
|---|---|---|
connected | Connection confirmation | Includes current status and language info |
result | Original text (STT) or real-time translation | Carries origin or translations depending on payload |
translation | Single-sentence re-translation result | Carries is_retranslation |
batch_retranslation | Batch re-translation result | When switching languages |
language_switch_start / language_switch_done | Batch re-translation progress | When switching languages in interpretation mode |
language_switched | Interpretation language switch | Interpretation mode |
translation_language_removed | A translation language was removed | After a successful multi-language switch_language with op=remove |
segment_discarded | Segment discarded | That sid will receive no further events; clear it from any "translating" state |
status | Status notification | Pause / resume / stop |
viewers | Current viewer count | When audience sharing is enabled |
subtitle_closed | Sharing closed, connection ended | Audience side; stop reconnecting on receipt |
speaker_renamed | Speaker renamed | Multi-speaker mode |
speaker_reassigned | Single-sentence speaker change | Multi-speaker mode |
speakers_merged / speakers_auto_merged | Speakers merged | Multi-speaker mode |
Each event's
datais identical to the corresponding message on the recording host's WebSocket; see WebSocket Events for full field definitions. Only the commonly used floating-subtitle events are listed below.
Event Formats
1. connected - Connection Confirmation
{
"status": "live",
"recognition_mode": "single",
"source_lang": "zh-TW",
"translation_languages": ["en-US", "ja-JP"]
}
| Field | Type | Description |
|---|---|---|
status | string | Recording status: live / paused / reconnecting / ended |
recognition_mode | string | Recognition mode: single / multi_speaker / multi_language |
source_lang | string | Source language (non-interpretation mode) |
conversation_languages | array | The two languages in interpretation mode (present only in interpretation mode, replaces source_lang) |
translation_languages | array | List of target translation languages (the current authoritative set: reflects mid-recording additions/removals, not the value at recording start). Omitted entirely when there is no target language |
2. result - Original Text (STT)
When data carries origin, it is original text:
{
"action": "result",
"origin": {
"sid": 1,
"language": "zh-TW",
"detected_language": "zh-TW",
"text": "大家好",
"is_final": true,
"speaker_id": "Guest-1",
"speaker_label": "Royx",
"start_time": "00:05"
}
}
| Field | Type | Description |
|---|---|---|
sid | number | Sentence ID |
language | string | Source language. In multi-language transcription it is determined per sentence and varies with the language actually spoken |
detected_language | string | Detected language of this sentence (for multi-language/interpretation positioning; in multi-speaker mode it is the fixed source language and must not be used as a language badge) |
text | string | Original text |
is_final | boolean | false = interim (will be replaced in place by later results with the same sid); true = finalized |
speaker_id | string | Original speaker ID (multi-speaker mode, optional) |
speaker_label | string | Display label (multi-speaker mode, optional) |
start_time | string | Sentence start time, format mm:ss |
3. result - Real-time Translation
When data carries translations, it is a translation (real-time translation is one message per language):
{
"action": "result",
"translations": {
"en-US": {
"sid": 1,
"text": "Hello everyone",
"is_final": true
}
}
}
| Field | Type | Description |
|---|---|---|
translations | object | Key is the target language code, value is the translation result |
translations.<lang>.sid | number | Corresponding original sentence ID |
translations.<lang>.text | string | Translation text |
translations.<lang>.is_final | boolean | Whether this is the final result |
Speaker inheritance: Translation messages do not carry
speaker_id/speaker_label. The client should match bysidto an already-received original line and reuse its speaker information.
4. translation / batch_retranslation - Re-translation
When the user triggers re-translation during recording, an updated translation is sent with the same sid; the client should replace in place:
{
"action": "translation",
"sid": 1,
"translations": {
"en-US": {
"sid": 1,
"text": "Hi everyone",
"is_final": true,
"is_retranslation": true
}
}
}
5. status - Status Notification
{
"action": "status",
"status": "paused",
"message": "Speech recognition paused"
}
| Field | Type | Description |
|---|---|---|
status | string | Machine-readable recording lifecycle state: live (resumed) / paused / ended (stopped). The client should act on this field: paused → freeze the display, ended → close the floating subtitle window, live → resume the display. |
message | string | Human-readable status text (format not guaranteed; do not parse it to determine state — always rely on the status field). |
When to close the floating subtitle window: the floating subtitle SSE stream does not close automatically when the recording stops (the connection is kept open by design). The host's own floating subtitle window must close itself on this event's
status: "ended", otherwise it will freeze on the last sentence. Thestatusfield appears only on thepause/resume/stoplifecycle transitions; otherstatusevents, such asset_nameandstart_speakingin manual conversation mode, do not carry it.
6. speaker_renamed / speaker_reassigned / speakers_merged - Speaker Events
Multi-speaker mode only; fields are identical to the corresponding events in WebSocket Events. On receipt, the client updates the display labels of the relevant sentences.
{
"action": "speaker_renamed",
"speaker_id": "Guest-1",
"new_label": "Royx",
"affected_sids": [1, 3, 5]
}
7. language_switched - Interpretation Language Switch
Interpretation mode only.
{
"action": "language_switched",
"active_lang": "en-US",
"translation_lang": "zh-TW"
}
8. viewers - Viewer Count
When audience sharing is enabled, the feed pushes the current viewer count; sent once on connection, then updated whenever the count changes.
{
"count": 3,
"max": 10
}
| Field | Type | Description |
|---|---|---|
count | number | Current number of audience members (excluding the owner) |
max | number | Audience limit (server-configured, default 10; rely on this field rather than hardcoding a value) |
9. subtitle_closed - Sharing Closed
When the host "closes sharing" or "stops the recording", the server proactively sends this event and then ends the audience connection (the audience side is a passive receiver).
{ "action": "subtitle_closed" }
On receipt, the audience client should show "sharing has ended" and stop reconnecting (sharing is now closed, so reconnection is rejected). The host's own connection is unaffected by this event.
Replay and Reconnection
- On connection, the most recent finalized original and translation lines are replayed first (for late joiners / reconnection).
- Speaker events are included in replay:
speaker_renamed/speaker_reassigned/speakers_merged/speakers_auto_mergedare replayed in original order (after the sentences they affect). Clients should handle them during replay exactly as in live mode — retroactively update the speaker labels of existing sentences byaffected_sids— so reconnecting or late-joining viewers do not see pre-rename speaker names. EventSourcereconnects automatically; as long as thefeed_tokenis still valid, finalized content from the interruption is replayed again after reconnection. Interim (is_final:false) drafts are not replayed and are naturally updated by subsequent results.
Heartbeat
The SSE connection uses a heartbeat to keep the connection alive:
- Interval: 15 seconds
- Format: SSE comment (starting with
:) - No client handling required; browsers ignore it automatically
: heartbeat
Frontend Example
async function connectSubtitle(taskId, apiKey, lang = null) {
// 1. Exchange for a feed_token
const res = await fetch(
`https://vas-poc.vurbo.ai/api/v1/auth/tasks/${taskId}/subtitle-feed-token`,
{ method: 'POST', headers: { 'X-API-Key': apiKey } }
);
if (res.status === 425) {
// Recording not ready yet; retry after a short delay
return;
}
const { token } = await res.json();
// 2. Connect to SSE
let url = `https://vas-poc.vurbo.ai/tasks/${taskId}/subtitle?feed_token=${token}`;
if (lang) url += `&lang=${lang}`;
const eventSource = new EventSource(url);
eventSource.addEventListener('connected', (e) => {
const data = JSON.parse(e.data);
console.log(`Status: ${data.status}, source: ${data.source_lang}`);
});
eventSource.addEventListener('result', (e) => {
const data = JSON.parse(e.data);
if (data.origin) {
// Original: replace in place by sid + is_final
console.log(`[${data.origin.sid}] ${data.origin.text}`);
} else if (data.translations) {
for (const [lang, t] of Object.entries(data.translations)) {
console.log(`Translation (${lang}): ${t.text}`);
}
}
});
eventSource.addEventListener('status', (e) => {
console.log(`Status: ${JSON.parse(e.data).message}`);
});
eventSource.onerror = () => {
// 410 = ended, 401 = token invalid; EventSource reconnects automatically, re-exchange the token if needed
};
return eventSource;
}
Version: V1.24.1 Last Updated: 2026-10-07