REST API

Imports API

Endpoint Overview

MethodEndpointDescription
POST/api/v1/imports/check-quotaCheck credit
POST/api/v1/importsUpload audio file
GET/api/v1/imports/{importId}Query import status
GET/api/v1/importsGet import list

POST /api/v1/imports/check-quota

Description

Check whether the user has enough credit to upload an audio file of a given duration. We recommend calling this API for a pre-check before uploading, so you avoid discovering that the credit is insufficient only after uploading a large file.

Authentication

Header: X-API-Key (see Authentication)

Request Parameters

ParameterLocationTypeRequiredDescription
duration_msbodyintegerYesAudio duration (milliseconds; 1 second to 10 hours by default -- the actual bounds follow the deployment setting and match the duration check applied after upload)

Request Example

curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports/check-quota" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
  -H "Content-Type: application/json" \
  -d '{"duration_ms": 3600000}'

Success Response

HTTP 200

{
  "data": {
    "allowed": true,
    "reason": null,
    "is_unlimited": false,
    "remain_quota": 480.0,
    "duration_minutes": 60,
    "estimated_points": 60.0
  }
}

Response Field Description

FieldTypeDescription
data.allowedbooleanWhether the upload is allowed (true when credit is sufficient or the plan allows it)
data.reasonstring | nullWhy the upload is not allowed: null (allowed) / insufficient_credit (insufficient credit; topping up resolves it) / plan_not_allowed (the plan does not include audio import; a plan upgrade is required) / plan_daily_limit_reached (today's plan usage is exhausted and resets tomorrow; topping up does not help, added in v1.16.4)
data.is_unlimitedbooleanWhether the account is unlimited (no point limit)
data.remain_quotafloat | nullRemaining credit points; null when unlimited
data.duration_minutesintegerEstimated audio duration (minutes, rounded up)
data.estimated_pointsfloatEstimated points to be charged (STT-based estimate; actual charge also includes translation/diarization, finalized at upload)

remain_quota semantics change (v1.9.0): for API Keys with a dedicated credit allotment, this returns the credit actually available to that key (its dedicated allotment, not the account's total balance); accounts without dedicated allotments see no change.

Specific Error Codes

This endpoint has no specific error codes; it may only return common authentication errors. Every rejection is returned as HTTP 200 with allowed: false and a reason.

Preflight reasonWhat the actual upload returns
plan_not_allowed403 plan_feature_not_allowed
insufficient_credit402 stt_quota_exceeded
plan_daily_limit_reached402 plan_daily_limit_reached (deliberately the same string, so you can map them 1:1)

Note: The preflight is a prediction, not a guarantee. Between an allowed: true response and the actual upload, other uploads or the per-minute settlement of an ongoing recording may still consume the quota, and the upload will then be rejected.


POST /api/v1/imports

Description

Upload an audio file for speech recognition and translation. After a successful upload, processing runs in the background, and you can track progress through the Query import status API.

Authentication

Header: X-API-Key (see Authentication)

Request Parameters (multipart/form-data)

ParameterTypeRequiredDescription
filefileYesAudio file (mp3/wav/m4a, max 500MB). The format is determined from the file's actual content; when the extension does not match the content (for example a .mp3 name with WAV content), the file is processed according to its content
transcription_languagesstringYesTranscription languages (JSON array string, e.g. '["zh-TW"]'). Up to 10, no duplicates
translation_languagesstringNoTranslation languages (JSON array string, e.g. '["en-US"]'). Up to 12, no duplicates
recognition_modestringYesRecognition mode: single (single speaker) / multi_speaker (multiple speakers). multi_language or multi_channel returns 422 import_recognition_mode_unsupported
summary_templatestringNoSummary template identifier (max 50 characters; must be an enabled summary category slug)
summary_modestringNoSummary mode: builtin (default, uses summary_template) or custom (uses summary_prompt). Omitted = uses summary_template
summary_promptstringNoFull custom prompt for custom mode (max 3000 chars, fully replaces the built-in template). Required for custom, prohibited otherwise
summary_prompt_slugstringNoCustom identifier for custom mode (max 64 chars, pass-through, not validated). Required for custom, prohibited otherwise
terminologystringNoTerminology list (JSON object string; see format below)
fuzzy_correctionstringNoFuzzy correction rules (JSON object string; see format below)
translation_dictstringNoTranslation dictionary (JSON object string; see format below)
callback_urlstringNoWebhook callback URL (notifies on completion/failure, max 2048 characters)

Webhook Notification: Once callback_url is set, you receive an import.completed event when the import finishes and an import.failed event when it fails. See the Webhook Guide.

Text Processing Parameter Formats

Terminology (terminology): Improves recognition accuracy for specific terms

{
  "zh-TW": [
    { "term": "Speaker Diarization" },
    { "term": "Real-time Transcription" }
  ]
}
  • Use the language code as the key and the term array as the value
  • term: Term text (required, max 100 characters)
  • Up to 500 terms per language, and no more than 500 across all languages combined
  • Up to 4000 fuzzy-correction rules per language, and no more than 4000 across all languages combined

The numbers above are defaults: the limit actually in force can be tuned per environment; always treat the 422 response message as authoritative.

Note: Both limits are enforced; exceeding either one returns 422 with the actual count. Plan multi-language vocabularies against the combined total.

Fuzzy Correction (fuzzy_correction): Corrects misspellings that sound different from the term

Usually no manual configuration is needed — misspellings that sound the same are already covered by terminology. It is needed only when the misspelling and the correct term sound different.

{
  "zh-TW": [
    { "correct": "Speaker Diarization", "incorrect": ["Speaker Diorization", "Speaker Diarizaion"] }
  ]
}
  • Use the language code as the key and the correction rule array as the value
  • correct: Correct term (required, max 200 characters)
  • incorrect: a list of incorrect variants (conditionally required, each up to 200 characters)

Supplying only the correct term: when correct is Chinese (contains Han characters), incorrect may be omitted entirely — the system matches by pronunciation, and spellings in the transcript that sound the same or nearly the same are corrected back to correct.

{ "fuzzy_correction": { "zh-TW": [{ "correct": "艾思通" }] } }

No misspellings need to be listed above: 愛思通, 愛時通, 愛司東 and 愛似通 are all corrected. Only spellings that sound quite different (愛自動, say) or have a different number of syllables (愛松) still need to be listed in incorrect.

Note: Both conditions must hold: the language must be Chinese (zh-TW, zh-CN, zh-HK and so on) and correct must contain Han characters. Otherwise incorrect remains required — omitting it in those cases would have no effect at all, and accepting it would leave you believing the setting took.

  • case_insensitive: Whether this rule's variants match regardless of case (optional, defaults to false = exact-case matching)

Case sensitivity: case_insensitive is optional and defaults to false (exact-case matching). When set to true, every incorrect variant in that rule matches regardless of case. The flag is per rule — the same correct term can be split across several rules with different settings, for example making variants that cannot collide with ordinary words case-insensitive while keeping variants that could hit a personal name exact. It has no effect on Chinese rules (Chinese has no letter case).

When the same incorrect variant appears in more than one rule: collisions are resolved on incorrect (the variant), not on correct. The case flag resolves to strict wins (if any rule leaves case_insensitive off, that variant is matched with exact case). Splitting one correct term across several rules is therefore safe, as long as their incorrect variants do not overlap.

Note: Enabling it widens the false-positive surface: if ivo is case-insensitive, the personal name Ivo is replaced too.

Translation Dictionary (translation_dict): Specifies how proper nouns are translated

{
  "en-US": [{ "source": "Speaker Diarization", "target": "Speaker Diarization" }]
}
  • Top-level key: the target language code
  • source: Source term (required, max 200 characters)
  • target: The required translation for this language (required, max 200 characters)
  • case_sensitive: Whether the entry applies only on an exact-case match (optional, defaults to false = case-insensitive)
  • Up to 3000 entries per language

The previous format is still supported: the earlier array-of-entries format is still accepted, with identical content and behavior.

Case-flag comparison: the case switches in fuzzy_correction and translation_dict have opposite field names, and their default value produces opposite behavior —

BlockFieldDefaultDefault behavior
fuzzy_correctioncase_insensitivefalseStrict (case-sensitive)
translation_dictcase_sensitivefalsePermissive (case-insensitive)

Both default to false, yet one means strict and the other means permissive. Do not share a single variable between them or mirror one onto the other — getting it wrong produces no error at all, only matching behavior opposite to what you intended.

Request Example

Basic Request

curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
  -F "file=@meeting.mp3" \
  -F 'transcription_languages=["zh-TW"]' \
  -F 'translation_languages=["en-US"]' \
  -F "recognition_mode=multi_speaker"

Request with Text Processing Settings

curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
  -F "file=@meeting.mp3" \
  -F 'transcription_languages=["zh-TW"]' \
  -F 'translation_languages=["en-US"]' \
  -F "recognition_mode=multi_speaker" \
  -F 'terminology={"zh-TW": [{"term": "Speaker Diarization"}]}' \
  -F 'fuzzy_correction={"zh-TW": [{"correct": "Speaker Diarization", "incorrect": ["Speaker Diorization"]}]}' \
  -F 'translation_dict={"en-US": [{"source": "Speaker Diarization", "target": "Speaker Diarization"}]}'

Request with Webhook Callback

curl -X POST "https://vas-poc.vurbo.ai/api/v1/imports" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
  -F "file=@meeting.mp3" \
  -F 'transcription_languages=["zh-TW"]' \
  -F 'translation_languages=["en-US"]' \
  -F "recognition_mode=multi_speaker" \
  -F "callback_url=https://your-server.com/webhooks/vas"

Success Response

HTTP 202

{
  "data": {
    "import_id": "550e8400-e29b-41d4-a716-446655440000",
    "status": "pending",
    "stage": null,
    "progress": 0,
    "message": null,
    "original_filename": "meeting.mp3",
    "file_size": "15.2 MB",
    "task_id": "660f9500-f30c-52e5-b827-557766550000",
    "error_code": null,
    "error_message": null,
    "created_at": "2026-02-23T10:00:00.000Z",
    "updated_at": "2026-02-23T10:00:00.000Z",
    "downgraded_features": []
  }
}

Response Field Description

FieldTypeDescription
data.import_idstringImport ID (UUID)
data.statusstringStatus: pending
data.stagestring | nullProcessing stage (initially null)
data.progressintegerProgress percentage (initially 0)
data.messagestring | nullProcessing message
data.original_filenamestringOriginal file name
data.file_sizestringFile size (formatted)
data.task_idstringTask ID (populated as soon as the upload succeeds; you can navigate to the task immediately)
data.error_codestring | nullError code (populated only on failure)
data.error_messagestring | nullError message (populated only on failure)
data.created_atstringCreation time (ISO 8601)
data.updated_atstringLast update time (ISO 8601)
data.downgraded_featuresarraySub-features skipped due to downgrade (v1.9.0): when the unlimited plan includes audio import but not some sub-features (such as speaker diarization speaker_diarization or translation translation), those sub-features are skipped and the import proceeds as usual; empty array = nothing downgraded

Specific Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
import_file_too_large413File size exceeds the 500MB limitCompress or split the file
import_invalid_format415Unsupported audio formatUse mp3/wav/m4a format
import_recognition_mode_unsupported422This recognition mode is not supported for file imports (multi_language, multi_channel). data.details carries field: "recognition_mode" and supportedModes: ["single", "multi_speaker"]Use single or multi_speaker
stt_quota_exceeded402Available credits are insufficient for this import's estimated chargeTop up and upload again
plan_feature_not_allowed403The unlimited plan does not include audio importUpgrade the plan; query GET /api/v1/me/plan for the plan contents
plan_daily_limit_reached402The plan's daily usage limit has been reachedUpload again after the plan's reset (the next day)

GET /api/v1/imports/{importId}

Description

Query the processing status and progress of a specific import task.

Once the task created by an import is deleted, the corresponding import record is removed as well, and this request returns 404 import_not_found.

Authentication

Header: X-API-Key (see Authentication)

Request Parameters

ParameterLocationTypeRequiredDescription
importIdpathstringYesImport ID (UUID)

Request Example

curl -X GET "https://vas-poc.vurbo.ai/api/v1/imports/550e8400-e29b-41d4-a716-446655440000" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"

Success Response

HTTP 200

{
  "data": {
    "import_id": "550e8400-e29b-41d4-a716-446655440000",
    "status": "processing",
    "stage": "transcribing",
    "progress": 45,
    "message": "Recognizing speech...",
    "original_filename": "meeting.mp3",
    "file_size": "15.2 MB",
    "task_id": "660f9500-f30c-52e5-b827-557766550000",
    "error_code": null,
    "error_message": null,
    "created_at": "2026-02-23T10:00:00.000Z",
    "updated_at": "2026-02-23T10:05:00.000Z"
  }
}

Response Field Description

FieldTypeDescription
data.statusstringStatus: pending / processing / completed / failed
data.stagestring | nullProcessing stage: converting / transcribing / translating / summarizing
data.progressintegerProgress percentage (0-100)
data.task_idstring | nullTask ID (can be used with the Tasks API). Populated from a successful upload onward; you do not need to wait for processing to finish
data.error_codestring | nullError code on failure
data.error_messagestring | nullError message on failure (a general description without internal details)

The remaining fields are the same as the POST /api/v1/imports response.

status Transitions

pending → processing → completed
                    → failed
  • completed and failed are final states and never change afterward: an import marked failed never becomes completed and is not charged.
  • Temporary errors during processing are retried automatically, and the status stays processing while retrying; only when retries are exhausted is the import marked failed, and the import.failed Webhook is sent only once.
  • An import that does not start processing for a long time (stays pending) is marked failed with error_code PROCESSING_TIMEOUT.
  • error_message is a general description based on the error code, without internal details; provide the import_id when you need troubleshooting. See Error Codes for the list of codes.

stage Processing Stages

StageDescription
convertingConverting audio format
transcribingRecognizing speech
translatingTranslating
summarizingGenerating summary

Edge Case for the completed Status (v1.3.5)

When an audio file produces no recognizable speech content—due to silence, low volume, noise, or a recognition language that does not match the audio—the system still finishes with completed (not failed), and task_id is generated normally, but the transcript entries for the corresponding task are an empty array. The client should load the data via GET /api/v1/sse/history/transcribe/{taskId} and then decide whether to show an empty state based on the sentence count. See File Import Guide – Behavior When Audio Cannot Be Recognized.

Specific Error Codes

Error CodeHTTP StatusDescriptionRecommended Action
import_not_found404Import task not foundVerify that importId is correct; if the task created by the import has been deleted, the import record is removed as well

GET /api/v1/imports

Description

Get the user's list of import tasks (paginated). Import records whose tasks have been deleted do not appear in the list.

Authentication

Header: X-API-Key (see Authentication)

Request Parameters

ParameterLocationTypeRequiredDescription
per_pagequeryintegerNoItems per page (default 20)

Request Example

curl -X GET "https://vas-poc.vurbo.ai/api/v1/imports?per_page=20" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW"

Success Response

HTTP 200

{
  "data": [
    {
      "import_id": "550e8400-e29b-41d4-a716-446655440000",
      "status": "completed",
      "stage": null,
      "progress": 100,
      "message": null,
      "original_filename": "meeting.mp3",
      "file_size": "15.2 MB",
      "task_id": "660f9500-f30c-52e5-b827-557766550000",
      "error_code": null,
      "error_message": null,
      "created_at": "2026-02-23T10:00:00.000Z",
      "updated_at": "2026-02-23T10:15:00.000Z"
    }
  ],
  "meta": {
    "current_page": 1,
    "last_page": 3,
    "per_page": 20,
    "total": 55
  }
}

Response Field Description

FieldTypeDescription
dataarrayList of import tasks (each field is the same as the Query import status response)
meta.current_pageintegerCurrent page number
meta.last_pageintegerLast page number
meta.per_pageintegerItems per page
meta.totalintegerTotal number of items

Specific Error Codes

This endpoint has no specific error codes; it may only return common authentication errors.


Version: V1.24.1 Last Updated: 2026-09-28

Copyright © 2026