Pricing
Overview
Vurbo.ai uses a credit-based billing model: you are charged in credits based on the features and duration you actually use. Credits can be pre-loaded and are consumed on a first-in-first-out (FIFO) basis by expiry date.
- Real-time recording / broadcast: billed per minute (rounded up to the next full minute); the per-minute rate is the sum of the credits for the enabled features.
- Value-added services (summary, re-translation): billed by text volume, based on the character count of the transcript.
- Broadcast: in addition to the host-side content-processing fee, a "cloud translation" fee is charged based on the number of translation languages and the audience scale.
When real-time recording is charged (broadcasts excluded)
- Credits are deducted at the start of each minute: the first minute is deducted when the recording starts, and
session_startedarrives once that deduction completes; each following minute is deducted as it begins, and nothing extra is deducted when the recording stops - When the available credits do not cover a full minute, that minute is not deducted: at the start of a recording you receive
auth_quota_exceeded; during a recording, the recording ends before the next minute begins and you receivestt_quota_exceeded - Turning features on or off, or adding or removing channels or translation languages during a recording, changes the rate from the next minute
- Each minute boundary has a 1-second grace period: a 60.5-second recording counts as 1 minute, and a 61.5-second recording counts as 2 minutes
- Billing continues while paused: pausing does not end the recording; to stop billing, end the recording (
stop) - The time spent waiting to resume after a disconnect is not billed: per-minute billing resumes once the session is resumed
- Time during which the audio cannot be decoded is not billed: minutes after an
audio_decode_failederror and before the audio is recognized again are not deducted (a minute that has already started is still billed); per-minute billing resumes afterward, so the billed minutes of a recording may be fewer than its duration - A recording that receives no usable audio at all is not billed: any credits already deducted are refunded automatically
- Long silence ends the recording automatically: when no speech is detected for a continuous period (15 minutes by default), the recording ends automatically so that a forgotten recording is not billed indefinitely; time spent paused does not count. See Automatic End After a Long Silence
Rates on this page are in credits; see "Credit Pricing" below for the credit-to-currency conversion.
In addition to the credit model, unlimited plans — a contract-authorized fixed feature bundle — are also available; see Unlimited Plans below.
Credit Pricing
| Item | TWD | USD |
|---|---|---|
| 1 credit | 1.5 | ≈ 0.047 |
The USD amount is converted at a reference rate of USD 1 ≈ TWD 32 (TWD 1.5 ÷ 32 ≈ USD 0.047), for reference only; the actual value is converted at the current spot exchange rate at the time of the transaction.
Base Usage (Real-time Recording, credits / minute)
The per-minute rate for real-time recording is the sum of the credits for the enabled features below. "Base speech recognition" is always charged; the rest stack based on what is actually enabled.
| Feature | Credits / min | Notes |
|---|---|---|
| Base speech recognition | 1.0 | Always charged |
| Multi-language detection | 0.5 | Bilingual / multi-language source detection |
| Speaker identification | 0.5 | Speaker diarization |
| Sentence translation | 0.2 | Either this or real-time translation |
| Real-time translation | 0.5 | Either this or sentence translation |
| Translation from the 2nd language (sentence translation, each +1) | +0.2 | Per additional translation language |
| Translation from the 2nd language (real-time translation, each +1) | +0.4 | Per additional translation language |
| Text-to-speech (TTS) | 1.0 | Real-time voice playback |
| Professional vocabulary | Free | Terminology / fuzzy correction / translation dictionary |
Professional vocabulary is not charged: terminology, fuzzy correction and translation dictionaries are all free features and do not affect the per-minute rate.
The number of source languages is not charged: when specifying multiple transcription languages (automatic language detection), the number of source languages does not affect the rate. Surcharges apply only to translation output languages.
Sentence vs. real-time translation: only one is charged (the higher rate applies). When translating into multiple languages, each language from the 2nd onward is surcharged per the table above.
Interpretation Mode
| Item | Credits / min | Notes |
|---|---|---|
| Real-time interpretation (conversation) | 1.5 | Integrated rate; includes speech recognition + translation |
| + Text-to-speech (TTS) | +1.0 | Added when voice playback is enabled |
The interpretation integrated rate does not include text-to-speech. With TTS enabled the rate is 1.5 + 1.0 = 2.5 credits per minute; without TTS it is 1.5 credits per minute. TTS can be toggled at any time during a session, and the rate adjusts from the next minute onward.
Multi-channel Speaker Diarization (credits / minute)
When a recording connects multiple physical microphones, multi-channel speaker diarization (recognition_mode: "multi_channel") can be enabled: speaker identity is determined by the channel, and you can choose to have each channel recognized independently (per_channel) or have channels take turns speaking and share one recognition stream (shared). Multi-channel is a form of "speaker identification (speaker diarization)" and applies to real-time recording; see the WebSocket reference for the operations.
Per-minute rate (N = the number of channels counted for that minute, up to 8):
| Item | Credits / min | Notes |
|---|---|---|
| Base speech recognition | 1.0 | Always charged (the 1st channel is included) |
| Speaker identification | 0.5 | Multi-channel is a form of speaker diarization; always charged |
| Multi-channel surcharge | (N − 1) × 1.0 | 1.0 credit per minute for each channel from the 2nd onward (in shared mode N is always 1, so no surcharge) |
That is, per minute = 1.0 + 0.5 + (N − 1) × 1.0 credits.
How N is counted
- Credits are deducted at the start of each minute, and N is the number of channels open at that moment.
- A channel added mid-minute via
add_channelis counted from the next minute; if it is disabled before the next deduction, or can no longer be received after recognition has started, it is counted once in the next deduction. If several channels are added and then disabled one after another within the same minute, the extra count is the highest number of them open at the same time. - After
remove_channeldisables a channel, the rate drops from the next minute (the current minute was already deducted when it began). A channel that was already deducted at the start of the minute is not billed again after it is disabled. - A channel whose audio cannot be received is not counted: if a channel's audio cannot be received at the moment a minute is settled (the channel produces no transcript during that period), that minute does not count this channel; once it recovers, it is counted again from the next minute.
- The first minute is settled when the recording starts, before any audio has been received, so it is based on the number of channels opened.
- A channel that is open but has never sent any audio is still counted — every channel is advised to keep sending audio (silence included) even when no one is speaking.
Examples
- 3 channels: 1.0 + 0.5 + (3 − 1) × 1.0 = 3.5 credits / minute
- 8 channels for 10 minutes: 1.0 + 0.5 + (8 − 1) × 1.0 = 8.5 credits/min; 10 minutes = 85 credits
sharedmode (channels take turns speaking and share one recognition stream): regardless of how many channels are open, 1.0 + 0.5 = 1.5 credits / minute;add_channel/remove_channeldo not affect the rate
Stacking with other billing axes
- Translation (sentence / real-time translation, incl. the per-language surcharge), meeting summaries, re-translation and other billable items are billed per their own rate tables as usual; they stack directly on top of the multi-channel surcharge without affecting it.
- Multi-channel does not support text-to-speech (TTS) or broadcast, so neither of those fees applies.
Multi-channel rules under unlimited plans
- Multi-channel requires the plan to include the feature: if the plan does not include it,
startreturnsplan_feature_not_allowed. - A plan can set a cap on the number of simultaneously recognized channels: exceeding the cap causes
start/add_channelto be rejected on the spot withplan_feature_not_allowed(details.fieldset tomax_stt_streams), rather than being accepted first and interrupted later.
Audio Import (credits / minute)
Uploading audio for offline transcription. Batch processing uses a more economical recognition service, so the base rate is lower than real-time recording.
| Item | Credits / min | Notes |
|---|---|---|
| Import base speech recognition | 0.3 | Batch-processing rate |
| Import feature surcharge | Per base-usage table | Speaker identification / sentence translation (incl. per-language surcharge) stack per the table above |
The base speech recognition rate for imports is 0.3 credits per minute; all other features (speaker identification, sentence translation and its per-language surcharge) stack using the same table as real-time recording.
Value-added Services (by text volume)
| Item | Rate | Billing unit |
|---|---|---|
| Meeting summary (first generation) | 0.1 credit | Per 1,000 transcript characters |
| Re-generate summary (change template) | 0.1 credit | Per 1,000 transcript characters |
| Ad-hoc summary (client-supplied content, added in v1.9.1) | 0.1 credit | Per 1,000 content characters |
| Summary translation (client-supplied content, added in v1.17.0) | 0.1 credit | Per 200 content characters |
| Re-translation (full text) | 0.1 credit | Per 200 characters |
| Re-translation (single sentence) | Free | — |
Definition of "character": the actual character count of the transcript (UTF-8 characters). Each Chinese, Japanese or Korean character counts as 1; each English letter counts as 1.
Rounding: any partial billing unit is charged as a full unit, and each action is charged at least 0.1 credit. Example: a summary of 12,000 characters = 1.2 credits; 12,050 characters = 1.3 credits.
Summaries are billed on the transcript characters sent to the model, not on the length of the generated summary.
Single-sentence re-translation is free: only full-text re-translation is billed by text volume.
No summary when credits are insufficient: when a real-time recording ends because the available credits ran out, or when the available credits at the end cannot cover the summary fee, no summary is generated and you receive
summary_error(summary_insufficient_credit); the transcript and audio are still saved.
Broadcast Billing
Broadcast billing consists of two independent billing lines; the total per-minute rate is the sum of both:
- Host side (content processing): billed per the "Base Usage" table above (speech recognition, translation, TTS, speaker identification, etc.), independent of audience size. Translation is always billed at the real-time translation rate, regardless of the
realtime_translationvalue sent instart. - Cloud translation (content delivery): the number of translation languages and the maximum audience size are looked up separately and added together.
- Billing starts at the moment the broadcast goes live: the standby phase is not billed; when a broadcast goes from standby to live, the first minute starts at go-live.
Broadcasts are not charged the "from the 2nd translation language" surcharge: multi-language costs for broadcasts are fully covered by the "translation languages" item below, so the host side does not double-count the per-language surcharge.
Cloud Translation Fee (credits / minute)
Cloud translation fee per minute = translation-language rate + maximum-audience rate.
By number of translation languages
| Translation languages | Credits / min |
|---|---|
| Translation not enabled | 0 |
| 1 | 2 |
| 2 | 6 |
| 3–4 | 14 |
| 5–8 | 30 |
| 9–12 | 62 |
By maximum audience size
| Maximum audience | Credits / min |
|---|---|
| ≤ 50 | 1 |
| ≤ 100 | 3 |
| ≤ 500 | 6 |
| ≤ 1,000 | 12 |
| ≤ 2,000 | 18 |
| ≤ 5,000 | 38 |
| ≤ 10,000 | 70 |
| ≤ 20,000 | 137 |
| ≤ 30,000 | 203 |
The audience count is the maximum audience size configured when the broadcast is created (fixed for the whole session, not the live concurrent count). The default maximum audience size per account is 5,000; contact sales to raise this limit.
Billing Examples
TWD / USD figures are reference conversions (1 credit = TWD 1.5 ≈ USD 0.047); USD is calculated at the current spot rate and is for reference only. Summary fees are based on transcript character count; each example states the assumed character count.
Example 1 — Real-time recording (transcribe): transcript only, no translation, 60 minutes (transcript ≈ 12,000 characters)
- Recording = base speech recognition 1.0 → 1.0 credit/min; 60 minutes = 60 credits
- Meeting summary = 12,000 ÷ 1,000 × 0.1 = 1.2 credits
- Total = 60 + 1.2 = 61.2 credits (≈ TWD 91.8 / USD 2.87)
Example 2 — Real-time recording (transcribe): sentence translation into 1 language, 120 minutes (transcript ≈ 25,000 characters)
- Recording = base speech recognition 1.0 + sentence translation 0.2 → 1.2 credits/min; 120 minutes = 144 credits
- Meeting summary = 25,000 ÷ 1,000 × 0.1 = 2.5 credits
- Total = 144 + 2.5 = 146.5 credits (≈ TWD 219.8 / USD 6.87)
Example 3 — Real-time recording (transcribe): sentence translation into 6 languages, 30 minutes (transcript ≈ 6,000 characters)
- Recording = base speech recognition 1.0 + sentence translation (0.2 + 0.2×5 = 1.0 from the 2nd language) → 2.2 credits/min; 30 minutes = 66 credits
- Meeting summary = 6,000 ÷ 1,000 × 0.1 = 0.6 credits
- Total = 66 + 0.6 = 66.6 credits (≈ TWD 99.9 / USD 3.12)
Example 4 — Interpretation (conversation): TTS enabled, 10 minutes (transcript ≈ 2,000 characters)
- Interpretation = integrated rate 1.5 + TTS 1.0 → 2.5 credits/min; 10 minutes = 25 credits
- Meeting summary = 2,000 ÷ 1,000 × 0.1 = 0.2 credits
- Total = 25 + 0.2 = 25.2 credits (≈ TWD 37.8 / USD 1.18)
- Without TTS: 1.5 credits/min → 15 + 0.2 = 15.2 credits
Example 5 — Audio import: sentence translation into 1 language, 30 minutes (transcript ≈ 6,000 characters)
- Import = base speech recognition 0.3 + sentence translation 0.2 → 0.5 credits/min; 30 minutes = 15 credits
- Meeting summary = 6,000 ÷ 1,000 × 0.1 = 0.6 credits
- Total = 15 + 0.6 = 15.6 credits (≈ TWD 23.4 / USD 0.73)
Example 6 — Broadcast: real-time translation into 3 languages, max 500 viewers, 180 minutes (transcript ≈ 38,000 characters)
- Host side = speech recognition 1.0 + real-time translation 0.5 → 1.5 credits/min (no per-language surcharge for broadcasts)
- Cloud translation = 3 languages (3–4 tier) 14 + audience ≤ 500 tier 6 = 20 credits/min
- Per minute = 1.5 + 20 = 21.5 credits/min; 180 minutes = 3,870 credits
- Meeting summary = 38,000 ÷ 1,000 × 0.1 = 3.8 credits
- Total = 3,870 + 3.8 = 3,873.8 credits (≈ TWD 5,810.7 / USD 181.58)
Example 7 — Real-time recording (transcribe) + speaker identification, 60 minutes (transcript ≈ 12,000 characters)
- Recording = base speech recognition 1.0 + speaker identification 0.5 → 1.5 credits/min; 60 minutes = 90 credits
- Meeting summary = 12,000 ÷ 1,000 × 0.1 = 1.2 credits
- Total = 90 + 1.2 = 91.2 credits (≈ TWD 136.8 / USD 4.28)
Example 8 — Audio import: transcript only, no translation, 45 minutes (transcript ≈ 9,000 characters)
- Import = base speech recognition 0.3 → 0.3 credits/min; 45 minutes = 13.5 credits
- Meeting summary = 9,000 ÷ 1,000 × 0.1 = 0.9 credits
- Total = 13.5 + 0.9 = 14.4 credits (≈ TWD 21.6 / USD 0.68)
Example 9 — Broadcast: real-time translation into 5 languages, max 1,000 viewers, 60 minutes (transcript ≈ 12,000 characters)
- Host side = speech recognition 1.0 + real-time translation 0.5 → 1.5 credits/min (no per-language surcharge for broadcasts)
- Cloud translation = 5 languages (5–8 tier) 30 + audience ≤ 1,000 tier 12 = 42 credits/min
- Per minute = 1.5 + 42 = 43.5 credits/min; 60 minutes = 2,610 credits
- Meeting summary = 12,000 ÷ 1,000 × 0.1 = 1.2 credits
- Total = 2,610 + 1.2 = 2,611.2 credits (≈ TWD 3,916.8 / USD 122.40)
Example 10 — Full-text re-translation: transcript of a 60-minute recording (≈ 12,000 characters)
- Full-text re-translation = 12,000 ÷ 200 × 0.1 = 6 credits
- Total = 6 credits (≈ TWD 9 / USD 0.28)
- Re-translating individual sentences only: free
Unlimited Plans
In addition to the credit model, unlimited plans are available (v1.9.0): a contract authorizes a fixed feature bundle and usage limits, and using the features included in the plan is not charged per minute during the authorized period. The plan's feature bundle and its limits are configured per contract; you can query the current plan's contents, usage, and when a restriction lifts at any time via GET /api/v1/me/plan.
Plan rules
- A plan is bound to a single API Key: one authorization applies to one API Key. Other API Keys under the same account are unaffected and continue to be billed in credits; covering multiple API Keys requires a separate authorization for each.
- Features the plan does not include are rejected outright: calling a feature the plan does not include returns
plan_feature_not_allowed(HTTP 403 over REST; an error message over WebSocket) and does not fall back to credit billing. See Error Code Reference — Plan and Usage Limit Errors. - Broadcasting is never included in unlimited plans: broadcasts are billed in credits (see "Broadcast Billing" above).
- Base features (included in every plan): base speech recognition, professional vocabulary, summary, and full-text re-translation.
Configurable plan limits
| Limit | Description |
|---|---|
| Simultaneously recognized languages | Exceeding the plan's cap rejects the start (too_many_languages, with details.max set to the plan's cap) |
| Single-recording length | The maximum length of one recording |
| Usage-hour thresholds | Once a threshold is reached, the active recording is periodically stopped (daily_limit_disconnect) — a new recording can start immediately; reaching the daily hard limit returns daily_limit_reached (the REST pre-check gate returns plan_daily_limit_reached, HTTP 402), and usage resumes after the plan's reset (daily limits reset the next day) |
| Concurrent recordings | The number of simultaneous recordings per API Key; at the limit, concurrency_limit_reached is returned. The slot is released when the server sends task_complete, so you can start the next recording once you receive it |
| Multi-channel channel cap | The cap on simultaneously recognized channels for multi-channel speaker diarization; exceeding it rejects start / add_channel on the spot (plan_feature_not_allowed, with details.field set to max_stt_streams) |
When limits are checked: for real-time recording, the single-recording limit, the usage-hour thresholds, and the daily hard limit are all checked before each minute begins; once one is reached, the next minute does not start and is not counted toward usage. When the same API Key runs several recordings at once, after the periodic-stop threshold is reached, each recording stops before its own next minute begins.
Restarting on the same connection: after a recording is stopped you can send
startimmediately, butsession_startedarrives only after the previous recording finishes processing (afterstatus: "ended"), which can take from a few seconds to a few tens of seconds depending on the summary length.
Version: V1.24.1 Last Updated: 2026-10-07