Appendix
Changelog
V1.24.1
2026-10-07
Fix: Viewer API Can Be Called from Any Domain
- When your viewer page was on your own domain, the browser blocked broadcast information lookups (
GET /api/v1/viewer/broadcasts/{token}) and password verification (POST /api/v1/viewer/broadcasts/{token}/verify). Starting with this version, both endpoints can be called directly from a web page on any domain. - Do not send credentials with the request (such as
credentials: 'include'); otherwise the browser blocks the response. When a 429 is returned, theRetry-Afterheader can be read. - The viewer caption stream already accepted connections from any domain and is unaffected; put
viewer_access_tokenin the URL query parameter. - See Viewer REST API — Calling from a Browser.
Documentation: Terminology Troubleshooting Example Corrected
- The "Terminology has no effect" row under Troubleshooting used
zh-CNregistered withzh-TWrecognition as its example, but the two are in the same language family and the terminology still takes effect. The example now uses languages from different families. - Added: a language code that is not recognized (e.g.
zh) also keeps terminology from taking effect;inactive_languagesandunknown_languagesinconfig_updatedlist these languages. - See Terminology Guide — Troubleshooting.
V1.24.0
2026-10-07
Behavior change: Broadcasts Always Use Real-Time Translation
- Broadcasts (
broadcast) always use real-time translation;realtime_translationinstartis treated astruewhether it is sent asfalseor not sent at all. Billing is unchanged and still uses the real-time translation rate. - Both the host and viewers first receive interim translations of a sentence (
is_final: false) and then the finalized version (is_final: true). Match them bysidandlanguageand replace what is displayed; to show only finalized results, skip events withis_final: false. A host whosestartsentrealtime_translationasfalseor omitted it now also receives interim results. - A viewer who joins partway through receives only the most recent few sentences of original text and finalized translations; interim results are not resent.
- See Broadcast Viewer SSE — Translation Result, Voice Translation Actions — Broadcast Mode Description, and Pricing — Broadcast Billing.
Behavior change: start Checks the Length of the Name and Summary Language
- If
nameinstartis longer than 60 characters after trimming leading and trailing whitespace, orsummary_languageis longer than 20 characters,invalid_parameteris returned and the recording does not start;details.fieldnames the field. - Previously,
startcould returnauth_service_error, or the recording could start but notask_completewas received after it ended. - When
summary_modeiscustom, asummary_promptorsummary_prompt_slugcontaining only whitespace counts as not provided:startreturnssummary_mode_field_mismatch;set_summarykeeps the previous setting and returnssummary_mode_field_mismatchonly if none was ever set. - See Voice Translation Actions — start and Error Codes — General Errors.
Behavior change: start Checks Fields That Accept Only Fixed Values
- If
options.speaking_speed,options.profanity_handling,conversation_mode, ortts_modeinstarthas a value outside the accepted list,invalid_parameteris returned and the recording does not start;details.fieldnames the field anddetails.valid_valueslists the accepted values. - Values are case-sensitive; for example,
"Async"fortts_modeis rejected. If a field is not sent or is an empty string, the default is used as before. - These fields are checked for every recording type: an invalid
conversation_modeon a recording that is not a conversation, or an invalidtts_modewith TTS turned off, is also rejected. - Previously, an invalid value was treated as the default, while the settings in
resume_okafter reconnecting showed the value originally sent. - When switching with
tts_modeduring a recording, a value other thansyncorasyncreturnsinvalid_data; the mode does not change and notts_mode_changedis sent. Previously,tts_mode_changedwas sent even though the mode did not actually change. - See Voice Translation Actions — start and Voice Translation Actions — tts_mode.
Behavior change: start_speaking Replies with status on Success
- In manual conversation mode, a successful
start_speakingnow returns astatusmessage; previously nothing was returned on success. - If it is called again while already speaking, the final result of the previous sentence is sent as usual, followed by the same
status; no error is returned. - The floating subtitle stream also receives this
statusmessage, without thestatusfield; clients that rely on thestatusfield to detect pause, resume, and stop are unaffected. - See Voice Translation Actions — start_speaking and Floating Subtitle SSE — status.
Behavior change: The Recording Completion Event May Arrive Twice
- In rare cases,
recording.completedfor the same recording is delivered twice, with a differentdelivery_ideach time, so checkingdelivery_idalone will not catch it. - Also check
data.task_idtogether witheventto decide whether it has already been processed. A repeated delivery does not deduct credits twice. - See Webhook Callback Guide — Idempotent Processing.
Fix: Manual Speaking and Speaker Language Changes in Conversations Without speakers
- When
startfor a conversation (conversation) omitsspeakers, users 1 and 2 map to the twotranscription_languagesin order, andstart_speakingandset_speaker_languagework normally; previously they returnedconversation_invalid_speaker. speakersis now marked optional and must still contain exactly 2 entries when provided;conversation_missing_speakersin the error code reference is marked as no longer returned.- For conversations started without
speakers, the settings inresume_okafter reconnecting now also includespeaker_language_map. - See Real-Time Voice Translation Guide — Conversation Mode and Voice Translation Actions — start.
Documentation: When Conversation Mode Translates
- Conversation mode translates only finalized sentences;
realtime_translationhas no effect in conversation mode, and the result is the same whether or not you send it. The field was removed from the conversation request examples. - See Core Concepts — Real-Time vs. Non-Real-Time Translation.
Documentation: When Terminology Settings Take Effect
- Updates to
terminology,fuzzy_correction, andtranslation_dictduring a recording take effect from the next utterance after they are sent; the utterance being recognized at that moment is not guaranteed to use them. The guide previously described this as the next recognition turn boundary or as immediate. - See Terminology Guide — When Settings Take Effect.
Documentation: Javanese and Wu Chinese Support Translation
jv-ID(Javanese) andwuu-CN(Wu Chinese) can be used as translation sources and targets, so all 145 languages support translation; the exception stating they did not support translation was removed.- See Supported Languages List — Language Overview.
Documentation: tts_stop Reply Text Corrected
- The
statustext returned whentts_stopsucceeds is corrected to "TTS stopped". - See Voice Translation Actions — tts_stop.
Documentation: Broadcast Channel Name and Summary Language Rules
- When creating or updating a broadcast,
summary_languagemust be a language code from the supported language list; otherwise the request returns 422validation_failed. - The channel name cannot be changed after creation; a
namefield in an update is ignored without an error. The recording name does not reuse the channel name either. - An update with no updatable field returns 200 and nothing changes; the REST API reference previously said at least one field must be provided.
- See Broadcast Feature Guide — Step 1: Create a Broadcast and Broadcast Feature Guide — Step 7: Update Settings Dynamically.
V1.23.0
2026-10-05
Behavior change: Rate Limit for Floating Subtitle Audience Token Requests
- Audience Token requests for floating subtitles are now limited to 30 per minute per recording, counting all viewers together, including requests with an invalid share link.
- Requests are no longer counted separately by source, and a source that sends invalid share links repeatedly is no longer temporarily blocked from requesting Tokens for other recordings.
- When the limit is exceeded, HTTP 429 (
too_many_requests) is still returned with aRetry-Afterheader. - See Subtitle Feed Token API — Audience Sharing.
Behavior change: Sending start While the Service Is Shutting Down
- A
startsent while the service is shutting down (for example, for a maintenance update) receivesservice_shutdown; no new recording is started, and the connection is then closed. - Recordings already in progress are not affected and can be finished as usual.
- After receiving it, reconnect later and send
startagain. - See Error Codes — Session Errors.
Documentation: Webhook Guide Corrections
- Corrected the retry count: each notification is sent at most 5 times (the first delivery plus 4 retries), at intervals of 10, 30, 90, and 270 seconds; the guide previously said 5 retries and 6 attempts in total.
- Added that 3xx and 4xx responses are also retried and that redirects are not followed; to have a notification that was marked as failed sent again, contact us.
- Corrected when a Webhook Secret is generated: you can choose to generate it when creating an API Key, and one is generated automatically when you set a URL for a key without a secret (secrets generated through batch settings are not shown in plaintext).
- Corrected the verification request sent when you save a URL: the system checks only whether the receiving endpoint responds with 2xx within 8 seconds and does not check the signature; if your endpoint verifies signatures, set the secret on it first.
- Corrected the repeat rule for credit balance notifications: each type is sent at most once every 24 hours per account and is re-evaluated after credits are added; also added that these two notifications are sent to every active API Key in the account that has a
webhook_url, one copy per key. - Added an "API Key for Each Notification" section: real-time recordings cannot specify a
callback_urland notify the API Key used to obtain the connection ticket; for broadcasts, the key that started the broadcast applies once the recording is finalized; no notifications are sent after an API Key is deleted. - Added
credit.lowandcredit.exhaustedto the event table in Core Concepts. - See the Webhook Callback Guide.
V1.22.1
2026-10-02
Fix: Numbers with Thousands Separators or Decimal Points in Real-Time Translation
- Numbers that contain thousands separators or decimal points (for example 1,200, 12,500, 1,200,000,000, and 3.14159) now stay intact in real-time translations and broadcast announcements, instead of keeping only their last part (for example, "1,200" translated as "200").
- These numbers keep the notation used in the original.
V1.22.0
2026-10-02
Behavior change: Translating Sentences That Mix Languages
- Words that are not in the translation language are always translated into it, even a single word.
- Names, companies, brands, products, places, all-caps abbreviations, model numbers, code, URLs, and email addresses are kept as in the original; words specified in the translation dictionary follow the dictionary.
- Parts of the original that are already in the translation language are left unchanged.
- Summary translation follows the same rule; numbers in summary translations keep the original's notation.
Behavior change: Other Language Columns in the Translation Dictionary Are Also Matched
- Words an entry has in other language columns are also matched against the original, so some words now follow the dictionary. For example, when the recognition language is Chinese, Japanese sentences are matched against the Japanese column; English words mixed into Chinese, Japanese, and similar sentences are matched against the English column.
- When translating into the recognition language (for example, Chinese to Chinese), sessions whose dictionary has other languages filled in also use the dictionary.
- See Terminology Guide — Other Language Columns Are Also Matched and Real-Time Voice Translation Guide — Translation Result.
V1.21.1
2026-10-01
Documentation: Response fields in multichannel shared mode
- Completed the description of multi-channel
sharedmode in responses. API behavior is unchanged:settings.channel_modeinresume_okcan beper_channelorshared.- In
sharedmode, thechannels[]entries insession_startedandresume_okdo not carrytranscription_languages; the session's languages are given bysettings.transcription_languages. - After resuming from a disconnect, each channel's
preparingstatus is reported in theresume_oksnapshot rather than by a separate event; when a channel starts producing text, it sends its ownchannel_statusevent (ready,reason: "reconnect").
- See Connection and Authentication — the settings object for field details.
V1.21.0
2026-10-01
New: Multi-Channel Shared Mode (Taking Turns)
Multi-channel adds channel_mode: "shared": the microphones take turns speaking and share one recognition stream, and the speaker is labeled by which channel the audio came in on. Suited to situations where only one person speaks at a time. See WebSocket API.
- Billing is always counted as 1 channel: 1.5 credits per minute (speech recognition 1.0 + speaker diarization 0.5); adding or removing channels does not change the rate.
stt_stream_countinchannel_statusis always1. - Client requirements: send only one channel at a time, keep sending silence on the current channel when nobody is speaking, and use
pcmonly. Languages are shared across the session: channels do not carrytranscription_languages, and the language cannot be changed during the recording (returnschannel_language_not_allowed). - The first channel in
channels[]cannot be removed (returnschannel_remove_not_allowed). - The
preparing/ready/errorstatus of the other channels follows the first channel. - Retroactive transcription after resuming from a pause covers the whole last 60 seconds, for all channels together.
- Known limitation: when the gap between speakers is shorter than about 0.8 seconds, the words of the two people may be merged into one sentence labeled with only one speaker.
- Environments where
sharedis not enabled still returninvalid_channel_mode.
Documentation: Multi-Channel Audio File Length Cap
- Clarified: the audio saved for one multi-channel recording has a total cap (about 70 minutes with 8 channels). Once the cap is reached, the audio file and the recording duration stop at that point, while the transcript and credit charges continue as usual.
V1.20.0
2026-09-30
New: API Key Self-Service Endpoints
Three new read-only endpoints return only this API Key's own data and remain available when credits run out. See API Key Self-Service.
GET /api/v1/me/credit-lots: the credit lots that charges draw on, with remaining credit and expiry time.GET /api/v1/me/usage: charge history, one entry per recording / broadcast and per import / summary / retranslation, paginated.GET /api/v1/me/key: the API Key's name, expiry date, monthly credit limit and this month's spend, concurrency limit, webhook URL, and whether a source IP restriction is configured.
V1.19.0
2026-09-30
New: Query Specific Tasks in the Task List
GET /api/v1/tasksadds thetask_ids[]parameter (UUIDs, 1–100 entries) and returns only the listed tasks that belong to the current account. IDs that do not exist or do not belong to the current account are skipped.statusstill applies; without it, only completed tasks are returned. To query regardless of status, addstatus=all.- The response format is unchanged.
Fix: Task List Failing for Accounts with Many Tasks
GET /api/v1/tasksnow returns normally for accounts with a large number of tasks.
For earlier versions, see Changelog Archive.
Version: V1.24.1 Last Updated: 2026-10-07