WebSocket Connection and Authentication
Table of Contents
Connection Information
| Item | Value |
|---|---|
| Endpoint | wss://vas-poc.vurbo.ai/ws |
| Protocol | WebSocket |
| Data Format | JSON |
| Auth Method | Ticket (see below) |
Authentication Method
VAS WebSocket uses a Ticket mechanism for authentication, passing a one-time Ticket via Sec-WebSocket-Protocol. For details, see Authentication.
Step 1: Obtain a Ticket
Use your API Key to exchange for a one-time Ticket via the REST API:
POST /api/v1/auth/ticket
X-API-Key: vas_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Response:
{
"ticket": "aBcDeFgHiJkLmNoPqRsTuVwXyZ012345",
"expires_in": 60
}
| Field | Type | Description |
|---|---|---|
ticket | string | One-time Ticket (32 chars) |
expires_in | int | Validity period (seconds) |
Step 2: Connect to WebSocket Using the Ticket
Place the Ticket in Sec-WebSocket-Protocol using the format ticket.{TICKET_VALUE}:
// Native browser support
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);
ws.onopen = () => {
console.log('Connected! Protocol:', ws.protocol);
// Start using the WebSocket...
};
ws.onerror = (error) => {
console.error('Connection failed:', error);
};
Node.js example:
const WebSocket = require('ws');
const ws = new WebSocket('wss://vas-poc.vurbo.ai/ws', [`ticket.${ticket}`]);
Ticket Characteristics
| Characteristic | Description |
|---|---|
| Validity period | 60 seconds |
| Usage count | Single use only (deleted immediately after use) |
| Security | The API Key is never exposed in the WebSocket connection |
| Replay protection | Atomic operations ensure single use |
Ticket Error Codes
| Error Code | HTTP Status | Description |
|---|---|---|
ticket_invalid | 401 | Ticket invalid or expired |
ticket_expired | 401 | Ticket expired |
ticket_already_used | 401 | Ticket already used |
ticket_validation_failed | 500 | Ticket validation failed |
For the complete API specification, see Auth Ticket API.
Message Format
All messages use a unified nested structure:
{
"type": "service type",
"data": { ... }
}
Maximum Size of a Single Message
A single WebSocket message has a maximum size (1 MB by default, adjustable per environment). Exceeding it closes the connection outright (close code 1009) with no error message of any kind — the client only sees the connection drop for no apparent reason.
Ordinary use never comes close to this limit: the recommended audio frame is 100 milliseconds (about 4 KB), and even a full second at a time is only a few tens of KB.
Note: The one thing that can hit it is a large glossary. The three glossary blocks can be sent across several
configmessages — a block you leave out is untouched and keeps its previous value, so splitting the send is safe. If the connection drops immediately after a largeconfigwith no error message at all, this is the direction to investigate.
Service Types
| type | Description |
|---|---|
health | Heartbeat mechanism |
voice-translation | Voice translation service |
error | Error message |
Error Message Format
When an error occurs, the server returns a message with type: "error":
{
"type": "error",
"data": {
"error_code": "auth_invalid_api_key",
"severity": "fatal",
"message": "Invalid API key",
"context": "auth",
"request_id": "req_abc123xyz789",
"timestamp": "2026-01-15T10:30:45.123Z"
}
}
| Field | Type | Description |
|---|---|---|
error_code | string | Error code (for programmatic handling) |
severity | string | Severity: fatal / error / warning |
message | string | Human-readable error message |
context | string | Error source category |
request_id | string | Request tracking ID |
timestamp | string | Time the error occurred (ISO 8601) |
For the complete list of error codes, see Error Code Reference.
Heartbeat Mechanism (Health)
Description
Used to confirm whether the WebSocket connection is healthy. We recommend sending a ping every 30 seconds; if no pong is received, treat the connection as dropped and reconnect. While the previous recording is still being processed after it ends (for example, while its summary is being generated), pong can be delayed by a few seconds to a few tens of seconds; allow enough waiting time and do not treat the connection as dropped just because a pong has not arrived yet.
Use Cases
- Maintain long-lived connections
- Detect connection status
- Prevent connection timeouts
Request - Ping
{
"type": "health",
"data": {
"action": "ping"
}
}
Response - Pong
{
"type": "health",
"data": {
"action": "pong"
}
}
Session Resume
If a WebSocket connection is unexpectedly dropped due to network instability, the client may
reconnect with its resume_token within a grace period (default 45 seconds) to rejoin the
original recording session — keeping the same task_id and sentence ids (sid), with the
transcript timeline continuing from the breakpoint. No need to restart the whole recording.
Use cases: mobile network switching, brief outages in elevators/tunnels, Wi-Fi roaming, etc. Brief network jitter that TCP recovers from within seconds does not drop the connection; session resume targets cases where the connection is actually severed.
Flow
1. start succeeds → session_started contains resume_token + resume_grace_seconds (store them)
2. connection drops → within the grace period:
a. obtain a NEW single-use Ticket from the backend
b. reconnect with Sec-WebSocket-Protocol: ["ticket.<newTicket>", "resume.<resume_token>"]
3. success → server returns resume_ok (with server_last_sid) → start a fresh audio stream as after start
failure → server returns a resume_* error → obtain a new Ticket, then fall back to a fresh start (unless the user has ended the recording)
New session_started fields
{
"type": "voice-translation",
"data": {
"action": "session_started",
"session_id": "...",
"task_id": "...",
"resume_token": "<43 chars; store until this session ends>",
"resume_grace_seconds": 45,
"server_time": 1749550000000
}
}
server_timeis the server's current time in unix milliseconds. The client can compute its clock skew relative to the server from(server_time, client time when received)as a reference; however, the grace-period countdown must still be measured against the client's wall clock (the instant the connection drops, the server can no longer push messages, so the client must compute the remaining time from the moment of disconnect itself).
Reconnect handshake
On reconnect, Sec-WebSocket-Protocol carries both the ticket and the resume token
(ticket always first):
const ws = new WebSocket(url, [`ticket.${newTicket}`, `resume.${resumeToken}`]);
- Every reconnect must obtain a new Ticket (tickets are single-use).
- The server validates the Ticket first (re-authentication + quota), then verifies the
resume_tokenownership against theuser_id/api_key_idderived from the Ticket.
Success response - resume_ok
{
"type": "voice-translation",
"data": {
"action": "resume_ok",
"session_id": "...",
"task_id": "...",
"recording_type": "transcribe",
"recognition_mode": "single",
"server_last_sid": 42,
"server_last_offset_ms": 125000,
"server_recording_ms": 127500,
"is_paused": false,
"settings": {
"speaking_speed": "normal",
"profanity_handling": "mask",
"audio_format": "pcm",
"transcription_languages": ["zh-TW"],
"translation_languages": ["en-US"],
"realtime_translation": true,
"auto_summary": true,
"summary_plain_text": false,
"summary_template": "meeting",
"summary_language": "zh-TW",
"summary_mode": "builtin",
"name": "Product Meeting"
},
"message": "Resumed the original session"
}
}
After receiving resume_ok:
- Start sending an audio stream again, just like after
start. For WebM/Opus you must send a brand-new container (restart the encoder); do not continue the old container. For PCM, simply continue sending. server_last_sidis the server's current last sentence id. If your locally received maximumsidis lower, a few sentences in flight were lost at the moment of disconnect — the live view may show a brief gap, but the complete content will appear in the final transcript when the recording ends.server_last_offset_msis the transcript timeline position (in milliseconds) of the breakpoint, based on the amount of audio already processed (not real elapsed time; the audio timeline is frozen during the disconnect). The client can use it to splice post-reconnect content at the correct timeline position, aligned together withserver_last_sid.server_recording_msis the recording-head timestamp (in milliseconds, including silence) that the transcript timeline resumes from after reconnect. The client uses it to align its recording-second header to the same timeline the transcript uses; the difference fromserver_last_offset_msis the trailing audio (silence or not-yet-finalized speech) after the last finalized sentence before the disconnect (omitted when 0).is_pausedis the paused state the server considers authoritative after resuming. Iftrue, the user was paused before the disconnect — after reopening the audio stream the client should pause immediately, send no audio, and keep the pause UI, instead of resuming recording on its own;falseor omitted means recording normally. This lets both a network reconnect and a full-page refresh realign the paused state (after a refresh the local pause memory is lost, so the server is authoritative).settingsis the set of recording settings currently held by the session, used to reconcile the client's local state with the server's authoritative values — see Thesettingsobject below.
Note: Do not mix the two notions of time. The transcript timeline (
server_last_offset_ms) is based on the audio file and is frozen during a disconnect; the grace reconnect window (resume_grace_seconds) is real wall-clock time and keeps counting down during a disconnect. Deciding "whether reconnect is still possible" must use wall-clock time, never the audio-file time position.
The settings object (added in v1.6.4)
resume_ok returns the recording settings currently held by the session, letting the client reconcile its local state against the server's authoritative values — eliminating drift from edge cases such as pause/resume, session resume, and multiple tabs.
Three principles:
- Current values: reflects changes made during recording via
set_speaking_speed/config/set_name/set_tts, not a snapshot taken atstart. - API format: values are returned in the public API format (e.g.,
speaking_speedreturns the level string"normal", not an the underlying numeric value). recording_type/recognition_modeare already at the top level ofresume_okand are not duplicated.
| Field | Type | Description |
|---|---|---|
speaking_speed | string | Current speaking speed (very_slow / slow / normal / fast / very_fast; returns normal if unset) |
profanity_handling | string | Profanity handling (mask / remove / show; returns the default mask if unset) |
audio_format | string | The audio format from start (pcm / webm). Resumed connections must keep using the same format (the server decodes with the original format and does not renegotiate) |
transcription_languages | string[] | Source languages (in conversation mode, use speaker_language_map as the source of truth after mid-recording language changes) |
translation_languages | string[] | Translation target languages |
realtime_translation | boolean | Real-time translation switch (always present; false is meaningful) |
tts_enabled | boolean | TTS switch (present when conversation TTS or single-mode TTS is enabled) |
tts_mode | string | TTS mode sync / async (same as above) |
tts_config | object | TTS voice settings (conversation bilingual / broadcast multilingual / single-mode single language; { "lang": { "voice", "speaking_rate" } }) |
conversation_mode | string | Conversation dialog mode auto / manual (changeable mid-recording via switch_conversation_mode; conversation mode only) |
speaker_language_map | object | Conversation speaker language map { "1": "zh-TW", "2": "en-US" } (changeable mid-recording via set_speaker_language; conversation mode only, present whether or not speakers was provided at start) |
terminology | array | Terminology (current values accumulated via config; no language grouping) |
fuzzy_correction | array | Fuzzy correction [{ "correct", "incorrect": [], "case_insensitive" }] (same as above; case_insensitive is omitted when false) |
translation_dict | object | array | Translation dictionary. The format matches the one you last sent — send the language-grouped format ({ "language code": [{ "source", "target", "case_sensitive" }] }) and you get that back; send the previous array-of-entries format and you get that back. The case_sensitive field is omitted when false |
auto_summary | boolean | Whether to auto-generate a summary (always present; false is meaningful) |
summary_plain_text | boolean | Whether the summary is plain-text output (always present; false is meaningful) |
summary_template | string | Summary template slug |
summary_language | string | Summary output language |
summary_mode | string | builtin / custom |
summary_prompt | string | Mode-aware: custom = full prompt / builtin = supplementary instructions (present only when set) |
summary_prompt_slug | string | Present in custom mode only |
name | string | Current recording name (reflects set_name changes) |
channel_mode | string | Multi-channel only (recognition_mode: "multi_channel"): channel mode (per_channel / shared) |
channels | array | Multi-channel only: current snapshot of all channels (including removed ones); see below |
Empty-collection fields (e.g., no terminology configured) are omitted entirely; treat "field absent" as "not configured".
Multi-channel fields channel_mode / channels[] (added in v1.10.0)
For multi-channel recordings (recognition_mode: "multi_channel"), settings additionally
returns channel_mode and channels[]. Session resume does not re-validate the start
parameters — after reconnecting, this snapshot is the client's only way to confirm that the
server is still in multi-channel mode and that the channel-to-language bindings are unchanged;
use it to overwrite the local channel state. In shared mode, channels are not bound to a
language; the session's languages are given by settings.transcription_languages.
Fields for each entry in channels[]:
| Field | Type | Description |
|---|---|---|
channel_id | int | Channel number |
speaker_name | string | Speaker name for this channel (present only when set) |
transcription_languages | string[] | Language bound to this channel (exactly one; current value after set_channel_language changes). Not present in shared mode |
status | string | Channel status: preparing (preparing, not yet producing text) / ready (has started producing text) / removed (disabled via remove_channel) / error (failed and cannot recover automatically). In shared mode, the other channels' preparing / ready / error follow the first channel |
Note: Removed channels also appear in the snapshot (
status: "removed") — this is not an omission: channel numbers are never reusable (including removed ones). After reconnecting, do not reuse aremovednumber inadd_channel— the server returnschannel_id_in_use. For the full multi-channel documentation, see Voice Translation Actions.
Failure responses
Resume failures always return severity: "error" (not fatal). Any failure code means the Ticket
attached to this handshake has been consumed; obtain a new Ticket before falling back to a
fresh start. If the user has already ended the recording (has sent stop), do not fall back, so that
you do not start a new recording the user did not ask for.
error_code | Meaning | Client action |
|---|---|---|
resume_token_invalid | Token invalid or not found; also returned after a service restart or update, because sessions waiting to be resumed are finalized at that point | New Ticket → fresh start |
resume_grace_expired | Grace period expired, session finalized | New Ticket → fresh start |
resume_ownership_mismatch | Token ownership mismatch (user/api_key) | New Ticket → fresh start |
resume_unavailable | Temporarily unavailable (this connection cannot resume the original session; also returned when the original session has been sent stop or has been ended and is still being processed) | New Ticket → fresh start |
When the service restarts or is updated, sessions waiting to be resumed are finalized right away: the recording is saved and the completion notification is sent as usual. Connections that are recording and not disconnected are unaffected; they first receive
service_shutdownand can finish the recording. Astartsent while the service is shutting down receivesservice_shutdown; no recording is started, and the connection is then closed. Reconnect later and start the recording again.
Content during the disconnect
- The few seconds of audio during the disconnect are not recovered (handled with "pause" semantics): the transcript timeline continues seamlessly, but the disconnected interval does not appear in the recording, and no extra charge applies for the disconnect. The minute in progress when the connection dropped was already deducted when it began and is not refunded; if the next minute begins during the disconnect, it is deducted once the resume succeeds.
- Conversation mode: a half-sentence already spoken but not yet finalized at the moment of disconnect is flushed as a provisional result.
- Single / speaker-diarization mode: a not-yet-finalized sentence at the moment of disconnect is not guaranteed to be preserved.
Trust model and security
api_keyis the trust boundary: the server uses a "last connection wins (takeover)" policy. Any client holding the sameapi_keyand the correspondingresume_tokencan take over the session within the grace period (the existing connection is silently closed). Therefore do not share a singleapi_keyacross trust domains; a B2B team sharing a key is treated as one trust boundary.- Transmitted only over WSS (TLS): both
resume_tokenandTicketare sent viaSec-WebSocket-Protocol; any plaintext (non-TLS) ingress is treated as credential leakage.
Reconnect implementation example (JavaScript)
The following wraps the full session-resume logic: store the token, detect disconnects, reconnect
within the grace period, and branch on resume_ok vs failure. Audio capture (startAudioStream,
etc.) depends on your format (PCM / WebM); for WebM you must reopen with a brand-new container
after resume_ok.
class ResumableVASClient {
// getTicket: async () => string (always obtain a NEW single-use Ticket per reconnect)
constructor({ wsUrl, getTicket, buildStartMessage }) {
this.wsUrl = wsUrl;
this.getTicket = getTicket;
this.buildStartMessage = buildStartMessage; // () => start message object
this.resumeToken = null;
this.graceSeconds = 45;
this.resumeDeadline = 0; // grace deadline (epoch ms); 0 = not resuming
this.closedByUser = false;
}
async start() {
this.closedByUser = false;
await this._connect(false);
}
stop() {
this.closedByUser = true;
this.resumeToken = null;
this.ws?.close(1000, 'client stop');
}
async _connect(isResume) {
const ticket = await this.getTicket(); // single-use; a new one on every reconnect
const protocols = isResume && this.resumeToken
? [`ticket.${ticket}`, `resume.${this.resumeToken}`]
: [`ticket.${ticket}`];
this.ws = new WebSocket(this.wsUrl, protocols);
this.ws.onopen = () => {
if (!isResume) this.ws.send(JSON.stringify(this.buildStartMessage())); // fresh start
// isResume: wait for resume_ok before reopening the audio stream (see onmessage)
};
// staleness guard: ignore late events from a stale socket (e.target !== this.ws) to avoid
// double connections after a resume fallback, or a stale onclose corrupting the new state.
this.ws.onmessage = (e) => { if (e.target === this.ws) this._onMessage(JSON.parse(e.data)); };
this.ws.onclose = (e) => { if (e.target === this.ws) this._onClose(); };
this.ws.onerror = () => {}; // onclose handles it uniformly
}
_onMessage(msg) {
const d = msg.data || {};
switch (d.action) {
case 'session_started':
this.resumeToken = d.resume_token; // store the token
this.graceSeconds = d.resume_grace_seconds || 45;
this.resumeDeadline = 0; // connection healthy; clear resume state
this.startAudioStream(); // start sending audio
break;
case 'resume_ok':
// resume succeeded: reopen the audio stream as after start (WebM must send a new container)
this.resumeDeadline = 0;
this.restartAudioStream();
break;
default:
if (msg.type === 'error' && String(d.error_code).startsWith('resume_')) {
// resume failed -> clear token, fall back to a fresh start (_connect obtains a new ticket)
this.resumeToken = null;
this.resumeDeadline = 0;
// the user has ended the recording (for example, stop() was called during a resume): do not start a new one
if (this.closedByUser) return;
this._connect(false).catch((err) => this.onFatal(err));
}
// other events (origin / translation, etc.) update the UI by sid
}
}
_onClose() {
if (this.closedByUser || !this.resumeToken) return this.onFatal?.();
// First disconnect: set the grace deadline
if (this.resumeDeadline === 0) {
this.resumeDeadline = Date.now() + this.graceSeconds * 1000;
}
if (Date.now() < this.resumeDeadline) {
// Within grace: attempt resume (retry with backoff while still within the window)
this._connect(true).catch(() => {
if (Date.now() < this.resumeDeadline) setTimeout(() => this._onClose(), 1000);
else this.onResumeExhausted?.();
});
} else {
this.onResumeExhausted?.(); // grace expired; give up resuming
}
}
// Implement these per your audio format:
startAudioStream() {} // begin capturing and sending audio frames
restartAudioStream() {} // reopen the audio stream (WebM = new container; PCM = continue)
onResumeExhausted() {} // could not resume within grace (treat as session end / notify user)
onFatal() {} // unrecoverable
}
Key points: (1) obtain a new Ticket on every reconnect; (2) reopen the audio stream only after
resume_ok; (3) on aresume_*error, fall back to a freshstart, but not if the user has ended the recording; (4) WebM must send a new container on reopen.
Version: V1.24.1 Last Updated: 2026-10-07