Terminology Guide
Table of Contents
- Overview
- Choosing Between the Three Blocks
- Selecting and Writing Terms
- Basic Configuration
- Configuration by Scenario
- When Settings Take Effect
- Important Notes
- Limits
- Validating a Glossary Before Saving
- Configuration Examples
- Troubleshooting
Overview
Terminology settings consist of three independent blocks that act at different stages of the recognition and translation pipeline:
| Block | Stage | Function | Determinism |
|---|---|---|---|
terminology | During recognition | Registers terms with the recognizer to improve recognition accuracy | Best-effort |
fuzzy_correction | After recognition | Replaces misrecognized words with the correct ones | Deterministic replacement |
translation_dict | During translation | Specifies how proper nouns are translated | Best-effort |
The three act at different stages and cannot substitute for one another. A term that must be both recognized correctly and translated consistently must be configured in terminology and translation_dict separately.
A single config request must provide at least one block; the three may be used individually or in combination.
How Terminology Participates in Homophone Correction
Once terminology is provided, the terms become the reference for homophone matching. Any span in the transcript that sounds the same or nearly the same but is written differently is corrected back to the spelling of the term:
| Term | Appears in transcript as | Corrected to |
|---|---|---|
| 紡拓會 | 訪拓會 | 紡拓會 |
| 語者分離 | 語這分離, 與者分離 | 語者分離 |
| 晶圓 | 晶園 | 晶圓 |
You do not need to list the possible misspellings in advance — matching is based on pronunciation, so coverage is not limited by how many variants you can think of.
Mixed Chinese-English terms: For a term like CVD製程, matching applies only to the Chinese portion; the Latin portion is left unchanged.
Applicable languages: Homophone matching applies to Chinese only (Traditional and Simplified are interchangeable, since they share pronunciation). Terms in Japanese, Korean or English do not participate in homophone matching — list their misspellings explicitly with fuzzy_correction.
Matching range: Both identical and near-identical pronunciations are covered, including accent differences such as
jinvsjing(final -n vs -ng) and retroflex vs non-retroflex initials. For the term 晶圓廠, for example, 金圓廠 in the transcript is corrected.Common-word protection: If the span in the transcript is itself an ordinary word (金元 or 反案, say), it is not changed even when it shares a pronunciation with a term — this prevents normal sentences from being altered. To force such a correction, list the misspelling explicitly with
fuzzy_correction: explicitly listed misspellings are not subject to common-word protection.
fuzzy_correctiongenerally does not require manual configuration. It is needed in three cases: the misspelling is itself an ordinary word and is blocked by common-word protection; the misspelling sounds very different from the correct term (a foreign brand name recognized as a phonetically unrelated word, say); or the language does not participate in homophone matching (Japanese, Korean, English).There is one further use that needs no misspellings at all: when Chinese terms exceed the 500-entry cap on
terminology, put the overflow intofuzzy_correctionwithcorrectonly — it is matched by pronunciation just the same (rule cap 4000).
When two terms share a pronunciation (registering both 公事包 and 公式包, say), the two registered words themselves are unaffected, but a third, unregistered homophonic spelling is assigned to one of them, and which one cannot be predicted. The Glossary Validation API can find cases like this before you save.
Conflict code:homophone_conflict
Choosing Between the Three Blocks
| Requirement | Recommended block |
|---|---|
| Company names, product names, or jargon must be recognized correctly | terminology |
| A term is consistently recognized as a known incorrect form and must be corrected | fuzzy_correction |
| The translation of a proper noun must be fixed | translation_dict |
| All of the above | All three, or terminology alone (see above) |
Selecting and Writing Terms
The effectiveness of terminology depends largely on term selection. The principles in this section apply to both terminology and fuzzy_correction.
Terms Suitable for Inclusion
| Type | Notes |
|---|---|
| Company and brand names | Invented words are readily broken into common homophonic characters |
| Product model numbers and project code names | Not everyday language; recognition lacks contextual basis |
| Speaker names | High proportion of uncommon characters |
| Domain-specific terminology | Jargon-dense scenarios such as medical, semiconductor, legal, and financial |
Prioritize terms that recur; those mentioned only once yield limited benefit.
Content Not Recommended for Inclusion
| Type | Reason |
|---|---|
| Common vocabulary | May cause otherwise correct content to be replaced with that term |
| Complete sentences | The unit of terminology is the term, not the sentence |
| Terms that may not be used | Consume quota, including those registered under languages not used in the session |
Terminology is not more accurate with more entries. Including excessive common vocabulary increases the likelihood of incorrect replacement.
Writing Terms
Terms are matched against spoken content rather than official full names.
| Situation | Recommended form |
|---|---|
| Only the short form is used in meetings | The short form |
| The product is referred to by a code in speech | The code |
| An English name has singular and plural forms | The form actually used, typically singular |
A term written more completely than what is spoken will not match.
For content written in Latin script (English, Vietnamese, Spanish, and so on), matching operates on whole words — a registered form is never applied to part of a longer word. For example, registering wafer has no effect on wafers in the transcript; register both forms if you need both. This rule applies only to Latin script; Chinese, Japanese, and Korean are unaffected.
Chinese terms are unaffected by casing; Traditional/Simplified differences are handled automatically and both forms need not be registered.
Rollout Recommendation
For initial configuration, include the 10–30 most critical terms, record an actual session, and revise based on the transcript: add terms still recognized incorrectly, and remove terms causing incorrect replacement (typically overly common vocabulary).
For recurring meetings or a single project, terminology content generally varies little and may be reused.
Basic Configuration
{
"type": "voice-translation",
"data": {
"action": "config",
"terminology": {
"zh-TW": [
{ "term": "語者分離" },
{ "term": "第三季營運計劃" }
],
"en-US": [
{ "term": "diarization" }
]
}
}
}
A config_updated event is returned on success.
Note: Sent only when the configuration is accepted. A rejected configuration returns
type: "error"instead, so your client must listen for that as well or the request appears to go unanswered.
Configuration by Scenario
All three scenarios use the same terminology format and rules; only the transport differs.
| Scenario | Transport | Value type |
|---|---|---|
| Live recording | WebSocket config action | JSON object |
| Broadcast | WebSocket config action (identical to live recording) | JSON object |
| File import | REST upload endpoint | JSON string |
Live Recording and Broadcast
The broadcast host side is an ordinary WebSocket connection; the config action has no additional restrictions or behavioral differences. Terminology configuration is not available to viewers.
Configuration may be sent before start or while recording is in progress.
File Import
Field names are identical to live recording, but values are JSON strings rather than objects:
terminology: '{"zh-TW":[{"term":"語者分離"}]}'
fuzzy_correction: '{"zh-TW":[{"correct":"語者分離","incorrect":["語這分離"]}]}'
translation_dict: '{"en-US":[{"source":"語者分離","target":"Speaker Diarization"}]}'
Note: Objects passed in the import scenario are rejected. See the file import endpoint reference for full specifications.
When Settings Take Effect
| Block | Before start | During recording |
|---|---|---|
terminology | Applied to the entire session | Supported; takes effect from the next utterance after it is sent; the utterance currently being recognized is not guaranteed |
fuzzy_correction | Applied to the entire session | Supported; takes effect from the next utterance after it is sent; the utterance currently being recognized is not guaranteed |
translation_dict | Applied to the entire session | Supported; takes effect from the next utterance after it is sent; the utterance currently being recognized is not guaranteed |
Updating any block during recording does not require reconnection, and the current recognition is not interrupted. When terminology is updated, the response also includes terminology_effective: "next_turn", indicating that new terms apply from the next utterance.
All three blocks are replaced wholesale. Resending
configreplaces rather than merges; adding a single term requires resending all existing terms.
Important Notes
Every item in this section corresponds to a situation the Glossary Validation API can detect before you save; each one ends with the matching Conflict code.
An Incorrect Variant That Is Also a Correct Term Breaks Normal Text
Where one string is at once the incorrect variant of rule A and the correct term of rule B (or a term in the terminology block), a user saying that word normally also has it rewritten into rule A's correct term.
{
"fuzzy_correction": {
"zh-TW": [
{ "correct": "報價", "incorrect": ["抱歉"] },
{ "correct": "抱歉" }
]
}
}
The example above turns 他說抱歉 into 他說報價. This is the only situation in this section that turns correct text into incorrect text, and it produces no error message at all.
Registering the same word as a term in the terminology block and as a correct term in fuzzy correction is not a problem — the two are different mechanisms (recognition weighting and after-the-fact correction), and using them together is good practice.
Conflict code:variant_shadows_term
The Two Case Flags Have Opposite Semantics
| Block | Field | Default | Default behavior |
|---|---|---|---|
fuzzy_correction | case_insensitive | false | Strict (case-sensitive) |
translation_dict | case_sensitive | false | Lenient (case-insensitive) |
Both default to false, but the resulting behaviors are opposite. Incorrect configuration produces no error message and yields only the inverse matching behavior. Do not share a single variable between them, and do not mirror one value to the other.
Where the same incorrect variant carries different case_insensitive settings across several rules, strict takes precedence.
Conflict code:case_flag_conflict
None of the Three Blocks Is Applied When Validation Fails
All three blocks are validated in one pass before any of them is applied. If validation fails for any one block, an error is returned and none of the three blocks is applied — the settings stay exactly as they were.
For example, sending a valid terminology block together with an oversized translation dictionary returns an error, and the terminology block does not take effect either. Correct the problem and resend the complete config (all three blocks are replaced wholesale; they are never merged with the previous settings).
Terminology Is Language-Scoped; Unused Entries Still Consume Quota
Terms and correction rules are keyed by language code and apply only to content matching that code. Rules registered under zh-TW do not act on English sentences.
Matching works at the level of the language family — codes sharing the part before the first hyphen belong to the same family. A rule registered under zh-CN therefore applies to zh-TW content as well, and vice versa; zh-TW and ja-JP do not affect each other.
An entry with an empty language code applies to every language. That is rarely the intent — make sure every entry carries a language code.
Entries registered under languages not used in the session still count toward the limits. Plan multilingual terminology against the combined total.
Registering the same term more than once under one language family likewise consumes quota for each copy (for example, registering it once under zh-TW and once under zh-CN). Correction behavior is unaffected, but terminology quota is spent for nothing.
Conflict code:duplicate_term
A Boost Outside the Valid Range Is Adjusted Automatically
A term's boost has a valid range of 0.5 to 5.0. When a value falls outside that range, the two paths behave differently:
| Path | When the value is out of range |
|---|---|
Live recording / broadcast (config) | Adjusted into range automatically, with no notice of any kind. Sending boost: 99 has exactly the same effect as sending 5.0, and nothing distinguishes the two |
| File import | Rejected outright (HTTP 422); no silent adjustment |
The same glossary producing different outcomes on the two paths is the easiest part of this to trip over.
Conflict code:boost_out_of_range(the response carriesvalueand theclamped_tovalue that actually applies)
Only One Rule Applies When an Incorrect Variant Is Duplicated
Collision is determined by incorrect rather than correct:
- Only one rule takes effect, and which one is not guaranteed (registration order included).
- The case flag resolves as strict takes precedence — if any rule leaves
case_insensitivedisabled, the variant is matched strictly.
Splitting a single correct term across multiple rules is safe provided their incorrect lists do not overlap. Where the data contains the same incorrect variant mapped to different correct terms, the later occurrence is ignored without notice — the Glossary Validation API can find duplicates like these before you send them.
Conflict code:variant_ambiguous
An Incorrect Variant Identical to Its Own Correct Term Has No Effect
Where incorrect contains a string identical to that rule's own correct value, the entry has no effect whatsoever (with case_insensitive enabled, differing only in casing counts as identical). It does not affect the result, but it does consume quota.
Conflict code:variant_equals_term
The Translation Dictionary Is a Hint, Not a Replacement
translation_dict guides translation wording through prompting and is best-effort. Compliance is not guaranteed for every sentence, and the proportion reliably honored decreases as entry count grows. This is inherent to prompt-based dictionaries and does not change if the limit is raised.
Use fuzzy_correction where deterministic replacement is required.
Under the same target language, when the same source is registered more than once only the last entry takes effect; the rest are ignored without notice.
Conflict code:dict_duplicate_source
Note: Dictionary entries are carried into translation requests individually, so entry count directly affects translation processing volume and cost. Only entries with a translation filled in for the target language are carried (when translating into the recognition language, entries with translations filled in for other languages are also carried); distributing translations across languages reduces the volume per request.
Other Language Columns Are Also Matched
source is written in the recognition language (the first transcription language). Besides source, the translations an entry has in other languages are also matched against the original text:
- When a sentence is not in the recognition language, the column for that sentence's language is matched. For example, with
source"會議",ja-JP"会議", anden-US"meeting": when "会議" in a Japanese sentence is translated into English, "meeting" is applied. - When a sentence is in a language that does not use the Latin alphabet, such as Chinese, Japanese, Korean, or Thai, the English column is also matched against English words mixed into the sentence.
- When translating into the recognition language, a matched word uses
sourceas its translation. So once the dictionary has other languages filled in, translations back into the recognition language also use the dictionary. - When a sentence's language cannot be determined (for example, with multi-language recognition), the columns for the session's other transcription languages are matched.
When matching from other language columns:
- English words are matched as whole words; for example,
maildoes not matchemail. - All-caps abbreviations (such as
AIandIT) are always case-sensitive. - Translations shorter than 2 characters are not matched.
- When the same word maps to different translations (for example, two entries both have "会議" in
ja-JPbut differenten-USvalues), that word is not used. - When the same word appears in more than one place, the sentence-language column takes precedence, then
source, then the English column.
Limits
| Block | Item | Default limit | Error code |
|---|---|---|---|
terminology | Number of terms (combined across all languages) | 500 | config_too_many_entries |
terminology | Length of a single term | 100 characters | config_term_too_long |
fuzzy_correction | Number of rules (combined across all languages) | 4000 | config_too_many_entries |
translation_dict | Number of entries (per language) | 3000 | config_too_many_dict_entries |
These are default values; effective limits may be adjusted per environment. Do not hard-code these numbers in an integration — the
detailsof a limit error carries bothcount(the number sent) andmax(the effective limit). Always readmax.
Validating a Glossary Before Saving
The situations under "Important Notes" above have one thing in common: none of them produces an error message. The glossary is accepted, the configuration succeeds, and the unexpected result only surfaces quietly during recording or translation.
The Glossary Validation API (POST /api/v1/glossary/validate) exists to point these out at the moment the user presses Save:
curl -X POST "https://<realtime-host>/api/v1/glossary/validate" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-H "Content-Type: application/json" \
-d '{
"fuzzy_correction": {
"zh-TW": [
{ "correct": "報價", "incorrect": ["抱歉"] },
{ "correct": "抱歉" }
]
}
}'
- Free: no points are deducted and no task is created
- All three glossary blocks are optional — whatever you send is what gets validated
- Every problem comes with the index of the entry, so a management UI can flag it directly
- The response carries no wording of its own; compose your own sentences from the codes
Recommended usage: have the glossary management UI call it once before saving. For live feedback while the user is still editing, send check_homophones: false to get a millisecond-scale response.
This endpoint validates the glossary itself. Passing validation does not guarantee that audio import will accept the same glossary — import applies stricter rules.
Configuration Examples
Company and Product Names
Configure terminology alone (misspellings that sound the same are corrected):
{ "terminology": { "zh-TW": [{ "term": "IPEVO" }] } }
Where the recognition error sounds different from the correct term (so homophone matching cannot cover it), add a rule:
{
"fuzzy_correction": {
"zh-TW": [{ "correct": "IPEVO", "incorrect": ["ltfo", "愛比"] }]
}
}
Casing for Brand Names
Where a Latin-script brand name may conflict with ordinary vocabulary, retain the default strict matching:
{
"fuzzy_correction": {
"en-US": [
{ "correct": "IPEVO", "incorrect": ["Ipevo", "ipevo"] }
]
}
}
Note: Enabling
case_insensitivewidens the scope of incorrect replacement. Ifivois set to ignore casing, the personal nameIvois also replaced.
Fixing the Translation of a Proper Noun
{
"translation_dict": {
"en-US": [{ "source": "語者分離", "target": "Speaker Diarization" }],
"ja-JP": [{ "source": "語者分離", "target": "話者分離" }]
}
}
Multilingual Meetings
Register each language's terms separately:
{
"terminology": {
"zh-TW": [{ "term": "第三季營運計劃" }],
"ja-JP": [{ "term": "予算" }],
"en-US": [{ "term": "quarterly report" }]
}
}
Terms across all languages count toward a combined limit.
Troubleshooting
| Symptom | Possible cause |
|---|---|
No response after sending config | A rejected configuration returns type: "error", not config_updated; a client waiting only for the latter hangs until its own timeout. Note also that an empty object {} does not count as "provided" |
| Terminology has no effect | The registered language is in a different language family from the recognition language (e.g. ja-JP registered while recognition uses zh-TW; zh-CN and zh-TW are in the same family and still match), or the language code is not recognized (e.g. zh). inactive_languages and unknown_languages in config_updated list these languages |
config sent during recording did not affect the current utterance | All three blocks take effect from the next utterance after they are sent; the utterance currently being recognized is not guaranteed to be covered. This is expected |
| An error was returned and nothing changed at all | Validation runs in one pass before anything is applied; if any block fails, none of the three is applied |
| Case matching behaves inversely to expectations | The two flags have opposite semantics; see Important Notes |
| A correction rule has no effect | Its incorrect variant duplicates another rule; only one of them applies, and which one is not guaranteed |
| Translations are inconsistent | The translation dictionary is best-effort; use fuzzy_correction where deterministic replacement is required |
| Terminology rejected during import | The three import fields are JSON strings, not objects |
| A normally spoken word is rewritten into a different one | That word is also the incorrect variant of another rule; see the first item under Important Notes |
| You want to find every situation above before saving | Use the Glossary Validation API |
For error code definitions see the error code reference; for the field-level specification of the config action see the WebSocket API documentation.
Version: V1.24.1 Last Updated: 2026-10-07