Guides

Terminology Guide

Table of Contents

  1. Overview
  2. Choosing Between the Three Blocks
  3. Selecting and Writing Terms
  4. Basic Configuration
  5. Configuration by Scenario
  6. When Settings Take Effect
  7. Important Notes
  8. Limits
  9. Validating a Glossary Before Saving
  10. Configuration Examples
  11. Troubleshooting

Overview

Terminology settings consist of three independent blocks that act at different stages of the recognition and translation pipeline:

BlockStageFunctionDeterminism
terminologyDuring recognitionRegisters terms with the recognizer to improve recognition accuracyBest-effort
fuzzy_correctionAfter recognitionReplaces misrecognized words with the correct onesDeterministic replacement
translation_dictDuring translationSpecifies how proper nouns are translatedBest-effort

The three act at different stages and cannot substitute for one another. A term that must be both recognized correctly and translated consistently must be configured in terminology and translation_dict separately.

A single config request must provide at least one block; the three may be used individually or in combination.

How Terminology Participates in Homophone Correction

Once terminology is provided, the terms become the reference for homophone matching. Any span in the transcript that sounds the same or nearly the same but is written differently is corrected back to the spelling of the term:

TermAppears in transcript asCorrected to
紡拓會訪拓會紡拓會
語者分離語這分離, 與者分離語者分離
晶圓晶園晶圓

You do not need to list the possible misspellings in advance — matching is based on pronunciation, so coverage is not limited by how many variants you can think of.

Mixed Chinese-English terms: For a term like CVD製程, matching applies only to the Chinese portion; the Latin portion is left unchanged.

Applicable languages: Homophone matching applies to Chinese only (Traditional and Simplified are interchangeable, since they share pronunciation). Terms in Japanese, Korean or English do not participate in homophone matching — list their misspellings explicitly with fuzzy_correction.

Matching range: Both identical and near-identical pronunciations are covered, including accent differences such as jin vs jing (final -n vs -ng) and retroflex vs non-retroflex initials. For the term 晶圓廠, for example, 金圓廠 in the transcript is corrected.

Common-word protection: If the span in the transcript is itself an ordinary word (金元 or 反案, say), it is not changed even when it shares a pronunciation with a term — this prevents normal sentences from being altered. To force such a correction, list the misspelling explicitly with fuzzy_correction: explicitly listed misspellings are not subject to common-word protection.

fuzzy_correction generally does not require manual configuration. It is needed in three cases: the misspelling is itself an ordinary word and is blocked by common-word protection; the misspelling sounds very different from the correct term (a foreign brand name recognized as a phonetically unrelated word, say); or the language does not participate in homophone matching (Japanese, Korean, English).

There is one further use that needs no misspellings at all: when Chinese terms exceed the 500-entry cap on terminology, put the overflow into fuzzy_correction with correct only — it is matched by pronunciation just the same (rule cap 4000).

When two terms share a pronunciation (registering both 公事包 and 公式包, say), the two registered words themselves are unaffected, but a third, unregistered homophonic spelling is assigned to one of them, and which one cannot be predicted. The Glossary Validation API can find cases like this before you save.

Conflict code: homophone_conflict


Choosing Between the Three Blocks

RequirementRecommended block
Company names, product names, or jargon must be recognized correctlyterminology
A term is consistently recognized as a known incorrect form and must be correctedfuzzy_correction
The translation of a proper noun must be fixedtranslation_dict
All of the aboveAll three, or terminology alone (see above)

Selecting and Writing Terms

The effectiveness of terminology depends largely on term selection. The principles in this section apply to both terminology and fuzzy_correction.

Terms Suitable for Inclusion

TypeNotes
Company and brand namesInvented words are readily broken into common homophonic characters
Product model numbers and project code namesNot everyday language; recognition lacks contextual basis
Speaker namesHigh proportion of uncommon characters
Domain-specific terminologyJargon-dense scenarios such as medical, semiconductor, legal, and financial

Prioritize terms that recur; those mentioned only once yield limited benefit.

TypeReason
Common vocabularyMay cause otherwise correct content to be replaced with that term
Complete sentencesThe unit of terminology is the term, not the sentence
Terms that may not be usedConsume quota, including those registered under languages not used in the session

Terminology is not more accurate with more entries. Including excessive common vocabulary increases the likelihood of incorrect replacement.

Writing Terms

Terms are matched against spoken content rather than official full names.

SituationRecommended form
Only the short form is used in meetingsThe short form
The product is referred to by a code in speechThe code
An English name has singular and plural formsThe form actually used, typically singular

A term written more completely than what is spoken will not match.

For content written in Latin script (English, Vietnamese, Spanish, and so on), matching operates on whole words — a registered form is never applied to part of a longer word. For example, registering wafer has no effect on wafers in the transcript; register both forms if you need both. This rule applies only to Latin script; Chinese, Japanese, and Korean are unaffected.

Chinese terms are unaffected by casing; Traditional/Simplified differences are handled automatically and both forms need not be registered.

Rollout Recommendation

For initial configuration, include the 10–30 most critical terms, record an actual session, and revise based on the transcript: add terms still recognized incorrectly, and remove terms causing incorrect replacement (typically overly common vocabulary).

For recurring meetings or a single project, terminology content generally varies little and may be reused.


Basic Configuration

{
  "type": "voice-translation",
  "data": {
    "action": "config",
    "terminology": {
      "zh-TW": [
        { "term": "語者分離" },
        { "term": "第三季營運計劃" }
      ],
      "en-US": [
        { "term": "diarization" }
      ]
    }
  }
}

A config_updated event is returned on success.

Note: Sent only when the configuration is accepted. A rejected configuration returns type: "error" instead, so your client must listen for that as well or the request appears to go unanswered.


Configuration by Scenario

All three scenarios use the same terminology format and rules; only the transport differs.

ScenarioTransportValue type
Live recordingWebSocket config actionJSON object
BroadcastWebSocket config action (identical to live recording)JSON object
File importREST upload endpointJSON string

Live Recording and Broadcast

The broadcast host side is an ordinary WebSocket connection; the config action has no additional restrictions or behavioral differences. Terminology configuration is not available to viewers.

Configuration may be sent before start or while recording is in progress.

File Import

Field names are identical to live recording, but values are JSON strings rather than objects:

terminology:      '{"zh-TW":[{"term":"語者分離"}]}'
fuzzy_correction: '{"zh-TW":[{"correct":"語者分離","incorrect":["語這分離"]}]}'
translation_dict: '{"en-US":[{"source":"語者分離","target":"Speaker Diarization"}]}'

Note: Objects passed in the import scenario are rejected. See the file import endpoint reference for full specifications.


When Settings Take Effect

BlockBefore startDuring recording
terminologyApplied to the entire sessionSupported; takes effect from the next utterance after it is sent; the utterance currently being recognized is not guaranteed
fuzzy_correctionApplied to the entire sessionSupported; takes effect from the next utterance after it is sent; the utterance currently being recognized is not guaranteed
translation_dictApplied to the entire sessionSupported; takes effect from the next utterance after it is sent; the utterance currently being recognized is not guaranteed

Updating any block during recording does not require reconnection, and the current recognition is not interrupted. When terminology is updated, the response also includes terminology_effective: "next_turn", indicating that new terms apply from the next utterance.

All three blocks are replaced wholesale. Resending config replaces rather than merges; adding a single term requires resending all existing terms.


Important Notes

Every item in this section corresponds to a situation the Glossary Validation API can detect before you save; each one ends with the matching Conflict code.

An Incorrect Variant That Is Also a Correct Term Breaks Normal Text

Where one string is at once the incorrect variant of rule A and the correct term of rule B (or a term in the terminology block), a user saying that word normally also has it rewritten into rule A's correct term.

{
  "fuzzy_correction": {
    "zh-TW": [
      { "correct": "報價", "incorrect": ["抱歉"] },
      { "correct": "抱歉" }
    ]
  }
}

The example above turns 他說抱歉 into 他說報價. This is the only situation in this section that turns correct text into incorrect text, and it produces no error message at all.

Registering the same word as a term in the terminology block and as a correct term in fuzzy correction is not a problem — the two are different mechanisms (recognition weighting and after-the-fact correction), and using them together is good practice.

Conflict code: variant_shadows_term

The Two Case Flags Have Opposite Semantics

BlockFieldDefaultDefault behavior
fuzzy_correctioncase_insensitivefalseStrict (case-sensitive)
translation_dictcase_sensitivefalseLenient (case-insensitive)

Both default to false, but the resulting behaviors are opposite. Incorrect configuration produces no error message and yields only the inverse matching behavior. Do not share a single variable between them, and do not mirror one value to the other.

Where the same incorrect variant carries different case_insensitive settings across several rules, strict takes precedence.

Conflict code: case_flag_conflict

None of the Three Blocks Is Applied When Validation Fails

All three blocks are validated in one pass before any of them is applied. If validation fails for any one block, an error is returned and none of the three blocks is applied — the settings stay exactly as they were.

For example, sending a valid terminology block together with an oversized translation dictionary returns an error, and the terminology block does not take effect either. Correct the problem and resend the complete config (all three blocks are replaced wholesale; they are never merged with the previous settings).

Terminology Is Language-Scoped; Unused Entries Still Consume Quota

Terms and correction rules are keyed by language code and apply only to content matching that code. Rules registered under zh-TW do not act on English sentences.

Matching works at the level of the language family — codes sharing the part before the first hyphen belong to the same family. A rule registered under zh-CN therefore applies to zh-TW content as well, and vice versa; zh-TW and ja-JP do not affect each other.

An entry with an empty language code applies to every language. That is rarely the intent — make sure every entry carries a language code.

Entries registered under languages not used in the session still count toward the limits. Plan multilingual terminology against the combined total.

Registering the same term more than once under one language family likewise consumes quota for each copy (for example, registering it once under zh-TW and once under zh-CN). Correction behavior is unaffected, but terminology quota is spent for nothing.

Conflict code: duplicate_term

A Boost Outside the Valid Range Is Adjusted Automatically

A term's boost has a valid range of 0.5 to 5.0. When a value falls outside that range, the two paths behave differently:

PathWhen the value is out of range
Live recording / broadcast (config)Adjusted into range automatically, with no notice of any kind. Sending boost: 99 has exactly the same effect as sending 5.0, and nothing distinguishes the two
File importRejected outright (HTTP 422); no silent adjustment

The same glossary producing different outcomes on the two paths is the easiest part of this to trip over.

Conflict code: boost_out_of_range (the response carries value and the clamped_to value that actually applies)

Only One Rule Applies When an Incorrect Variant Is Duplicated

Collision is determined by incorrect rather than correct:

  • Only one rule takes effect, and which one is not guaranteed (registration order included).
  • The case flag resolves as strict takes precedence — if any rule leaves case_insensitive disabled, the variant is matched strictly.

Splitting a single correct term across multiple rules is safe provided their incorrect lists do not overlap. Where the data contains the same incorrect variant mapped to different correct terms, the later occurrence is ignored without notice — the Glossary Validation API can find duplicates like these before you send them.

Conflict code: variant_ambiguous

An Incorrect Variant Identical to Its Own Correct Term Has No Effect

Where incorrect contains a string identical to that rule's own correct value, the entry has no effect whatsoever (with case_insensitive enabled, differing only in casing counts as identical). It does not affect the result, but it does consume quota.

Conflict code: variant_equals_term

The Translation Dictionary Is a Hint, Not a Replacement

translation_dict guides translation wording through prompting and is best-effort. Compliance is not guaranteed for every sentence, and the proportion reliably honored decreases as entry count grows. This is inherent to prompt-based dictionaries and does not change if the limit is raised.

Use fuzzy_correction where deterministic replacement is required.

Under the same target language, when the same source is registered more than once only the last entry takes effect; the rest are ignored without notice.

Conflict code: dict_duplicate_source

Note: Dictionary entries are carried into translation requests individually, so entry count directly affects translation processing volume and cost. Only entries with a translation filled in for the target language are carried (when translating into the recognition language, entries with translations filled in for other languages are also carried); distributing translations across languages reduces the volume per request.

Other Language Columns Are Also Matched

source is written in the recognition language (the first transcription language). Besides source, the translations an entry has in other languages are also matched against the original text:

  • When a sentence is not in the recognition language, the column for that sentence's language is matched. For example, with source "會議", ja-JP "会議", and en-US "meeting": when "会議" in a Japanese sentence is translated into English, "meeting" is applied.
  • When a sentence is in a language that does not use the Latin alphabet, such as Chinese, Japanese, Korean, or Thai, the English column is also matched against English words mixed into the sentence.
  • When translating into the recognition language, a matched word uses source as its translation. So once the dictionary has other languages filled in, translations back into the recognition language also use the dictionary.
  • When a sentence's language cannot be determined (for example, with multi-language recognition), the columns for the session's other transcription languages are matched.

When matching from other language columns:

  • English words are matched as whole words; for example, mail does not match email.
  • All-caps abbreviations (such as AI and IT) are always case-sensitive.
  • Translations shorter than 2 characters are not matched.
  • When the same word maps to different translations (for example, two entries both have "会議" in ja-JP but different en-US values), that word is not used.
  • When the same word appears in more than one place, the sentence-language column takes precedence, then source, then the English column.

Limits

BlockItemDefault limitError code
terminologyNumber of terms (combined across all languages)500config_too_many_entries
terminologyLength of a single term100 charactersconfig_term_too_long
fuzzy_correctionNumber of rules (combined across all languages)4000config_too_many_entries
translation_dictNumber of entries (per language)3000config_too_many_dict_entries

These are default values; effective limits may be adjusted per environment. Do not hard-code these numbers in an integration — the details of a limit error carries both count (the number sent) and max (the effective limit). Always read max.


Validating a Glossary Before Saving

The situations under "Important Notes" above have one thing in common: none of them produces an error message. The glossary is accepted, the configuration succeeds, and the unexpected result only surfaces quietly during recording or translation.

The Glossary Validation API (POST /api/v1/glossary/validate) exists to point these out at the moment the user presses Save:

curl -X POST "https://<realtime-host>/api/v1/glossary/validate" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
  -H "Content-Type: application/json" \
  -d '{
    "fuzzy_correction": {
      "zh-TW": [
        { "correct": "報價", "incorrect": ["抱歉"] },
        { "correct": "抱歉" }
      ]
    }
  }'
  • Free: no points are deducted and no task is created
  • All three glossary blocks are optional — whatever you send is what gets validated
  • Every problem comes with the index of the entry, so a management UI can flag it directly
  • The response carries no wording of its own; compose your own sentences from the codes

Recommended usage: have the glossary management UI call it once before saving. For live feedback while the user is still editing, send check_homophones: false to get a millisecond-scale response.

This endpoint validates the glossary itself. Passing validation does not guarantee that audio import will accept the same glossary — import applies stricter rules.


Configuration Examples

Company and Product Names

Configure terminology alone (misspellings that sound the same are corrected):

{ "terminology": { "zh-TW": [{ "term": "IPEVO" }] } }

Where the recognition error sounds different from the correct term (so homophone matching cannot cover it), add a rule:

{
  "fuzzy_correction": {
    "zh-TW": [{ "correct": "IPEVO", "incorrect": ["ltfo", "愛比"] }]
  }
}

Casing for Brand Names

Where a Latin-script brand name may conflict with ordinary vocabulary, retain the default strict matching:

{
  "fuzzy_correction": {
    "en-US": [
      { "correct": "IPEVO", "incorrect": ["Ipevo", "ipevo"] }
    ]
  }
}

Note: Enabling case_insensitive widens the scope of incorrect replacement. If ivo is set to ignore casing, the personal name Ivo is also replaced.

Fixing the Translation of a Proper Noun

{
  "translation_dict": {
    "en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }],
    "ja-JP": [{ "source": "語者分離", "target": "話者分離" }]
  }
}

Multilingual Meetings

Register each language's terms separately:

{
  "terminology": {
    "zh-TW": [{ "term": "第三季營運計劃" }],
    "ja-JP": [{ "term": "予算" }],
    "en-US": [{ "term": "quarterly report" }]
  }
}

Terms across all languages count toward a combined limit.


Troubleshooting

SymptomPossible cause
No response after sending configA rejected configuration returns type: "error", not config_updated; a client waiting only for the latter hangs until its own timeout. Note also that an empty object {} does not count as "provided"
Terminology has no effectThe registered language is in a different language family from the recognition language (e.g. ja-JP registered while recognition uses zh-TW; zh-CN and zh-TW are in the same family and still match), or the language code is not recognized (e.g. zh). inactive_languages and unknown_languages in config_updated list these languages
config sent during recording did not affect the current utteranceAll three blocks take effect from the next utterance after they are sent; the utterance currently being recognized is not guaranteed to be covered. This is expected
An error was returned and nothing changed at allValidation runs in one pass before anything is applied; if any block fails, none of the three is applied
Case matching behaves inversely to expectationsThe two flags have opposite semantics; see Important Notes
A correction rule has no effectIts incorrect variant duplicates another rule; only one of them applies, and which one is not guaranteed
Translations are inconsistentThe translation dictionary is best-effort; use fuzzy_correction where deterministic replacement is required
Terminology rejected during importThe three import fields are JSON strings, not objects
A normally spoken word is rewritten into a different oneThat word is also the incorrect variant of another rule; see the first item under Important Notes
You want to find every situation above before savingUse the Glossary Validation API

For error code definitions see the error code reference; for the field-level specification of the config action see the WebSocket API documentation.


Version: V1.24.1 Last Updated: 2026-10-07

Copyright © 2026