Ad-hoc Summary SSE
Connection Info
| Item | Value |
|---|---|
| Base path | https://vas-poc.vurbo.ai/api/v1/sse |
| Protocol | HTTP + Server-Sent Events (SSE) |
| Data format | text/event-stream |
| Authentication | Header X-API-Key: {KEY} (this endpoint does not accept the key via query string) |
Note: This endpoint only accepts POST (JSON body), so the browser-native EventSource API cannot be used. Call it with the fetch API plus ReadableStream, or any HTTP client that can read streamed responses.
Endpoint Overview
| Method | Endpoint | Persists Result | Billed | Purpose |
|---|---|---|---|---|
| POST | /api/v1/sse/summary | No | Yes | Generate a summary from arbitrary text supplied in the request (not tied to any recording) |
Difference from Regenerate Summary: regeneration always uses the transcript stored on the server as input, while this endpoint uses the content supplied in the request — for content the server does not have, such as the full transcript of several merged recordings, or a transcript edited by the user. The result is never stored; it is only streamed back to the client.
Added in v1.9.1. An earlier endpoint once existed at the same path (removed in V1.8.0); this endpoint is a brand-new contract — authentication, billing, and duplicate-request rules all differ, so do not reuse old integration code.
Request Parameters (JSON body)
| Parameter | Type | Required | Constraints | Description |
|---|---|---|---|---|
content | string | Yes | ≤200,000 characters | The full text to summarize |
idempotency_key | string | Yes | ≤64 characters, A-Z a-z 0-9 . _ - | Duplicate-request identifier that prevents double billing (see "Duplicate Requests and Billing Guarantees" below) |
mode | string | Yes | "builtin" | "custom" | Summary mode: built-in template or custom prompt |
template | string | Required for builtin / forbidden for custom | A valid built-in template slug | Built-in template slug |
prompt | string | Required for custom / forbidden for builtin | ≤3000 characters | The customer's full prompt |
promptSlug | string | Required for custom / forbidden for builtin | ≤64 characters, Unicode, no control characters | Customer-defined identifier (returned as-is, not processed) |
language | string | No | - | Output language code for the summary (e.g. zh-TW, en-US); defaults to zh-TW when omitted |
plainText | boolean | No | Defaults to false | Request plain-text output (Markdown formatting symbols are removed automatically) |
Mutual-exclusion rules (same as Regenerate Summary):
- With
mode=builtin,promptandpromptSlugmust not be sent - With
mode=custom,templatemust not be sent, whilepromptandpromptSlugare required
Error Response Model (differs from other SSE endpoints — please note)
This endpoint is intended for backend integrations. Errors before the stream starts always return a real HTTP status code with a JSON body, instead of the "HTTP 200 + error event" convention used by other SSE endpoints:
| Phase | Response form |
|---|---|
| Authentication failure (401/403), validation failure (422), template missing or deactivated (404), all-whitespace content (400), insufficient credit (402), duplicate-request identifier conflict (409), rate limit exceeded (429) | Real HTTP status code + JSON error body |
| After the stream starts (generation failure, filtered content, etc.) | HTTP 200 + SSE error event |
Authentication only accepts the
X-API-Keyheader; API keys in the query string are not accepted.
The 422 JSON body carries per-field error messages under data.details.errors, so your program can act on them:
{
"type": "error",
"data": {
"error_code": "validation_failed",
"message": "Validation failed",
"details": { "errors": { "content": ["The content parameter is required"] } }
}
}
Billing
- Billed by the character count of
content: 0.1 credits per 1,000 characters (rounded up below 1,000; same rate as Meeting Summary and Regenerate Summary). - Billed only on successful generation; failed generation (including filtered content, or a stream that stalls or does not end normally) is not billed. A summary that is incomplete, with
truncatedindone, still counts as successful and is billed as usual. - Example: a 35,000-character
content= 3.5 credits; 1 character = 0.1 credits (minimum billing unit).
Duplicate Requests and Billing Guarantees
idempotency_key is supplied by the client (we recommend a stable identifier from your system, such as a merge-batch ID or revision ID). The system compares the entire request — content plus every summary parameter (mode / template / prompt / promptSlug / language / plainText) — to decide whether two calls are retries of the same piece of work:
| Scenario | Behavior |
|---|---|
| Retry with the same identifier + an identical request | Not billed again, but the summary is regenerated (the result may differ from last time; the previous result is not replayed); retries still work even when the account balance has run out (the original request was already billed) |
Same identifier + any field differs (content or any summary parameter) | Returns 409 summary_idempotency_key_conflict; nothing is generated or billed |
| Retry after a failed first request | Failure does not claim the identifier — retrying (even with different content) is treated as a brand-new request and billed normally |
Retries must carry exactly the same fields as the original request; when summarizing new content or switching prompts, always send a new
idempotency_key.
idempotency_keyis only valid within a single API key: the same identifier on different API keys is unrelated and billed separately. Re-sending the same piece of work through a backup key counts as two requests.Edge case: if the same identifier is used concurrently for two different requests (a client error), the first to complete claims it; the other is still billed and delivered normally, but any later retry with that identifier (regardless of content) returns 409 — switch to a new identifier.
Request Examples
builtin mode
curl -N -X POST "https://vas-poc.vurbo.ai/api/v1/sse/summary" \
-H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
-H "Content-Type: application/json" \
-d '{
"content": "(full transcript merged from multiple recordings…)",
"idempotency_key": "merge-batch-20260813-001",
"mode": "builtin",
"template": "meeting",
"language": "en-US",
"plainText": true
}'
custom mode
curl -N -X POST "https://vas-poc.vurbo.ai/api/v1/sse/summary" \
-H "X-API-Key: vas_..." \
-H "Content-Type: application/json" \
-d '{
"content": "(full transcript after user edits…)",
"idempotency_key": "revision-8f31c2",
"mode": "custom",
"prompt": "Extract the decisions and action items from the transcript…",
"promptSlug": "acme-meeting-v2",
"language": "en-US"
}'
Event Sequence
1. connected → connection confirmed
2. summary_regeneration → summary fragments (repeated N times, cumulative)
3. done → generation complete
or
3. error → generation failed (sse_summary_regeneration_failed); no done is sent and nothing is billed
connected
{ "message": "Ad-hoc summary stream connected (mode: builtin)" }
summary_regeneration
{ "text": "The meeting covered the following topics:\n1. Product development progress", "is_final": false }
| Field | Type | Description |
|---|---|---|
text | string | Cumulative summary content (with plainText=true, the is_final=true text is the cleaned plain text) |
is_final | boolean | Whether this is the final result |
done
{
"tokens_used": 123,
"final_content": "The meeting... (full summary content)",
"mode": "custom",
"template": "acme-meeting-v2",
"plain_text": true,
"characters_billed": 35000,
"charged": "3.5",
"idempotency_key": "revision-8f31c2",
"billed": true,
"summary_language": "zh-TW",
"prompt_snapshot": "Extract the decisions and action items from the transcript…"
}
| Field | Type | Description |
|---|---|---|
tokens_used | number | Total token usage |
final_content | string | Full summary content (cleaned plain text when plainText=true) |
mode | string | Summary mode: "builtin" or "custom" |
template | string | The identifier actually used — builtin → built-in template slug; custom → customer-defined identifier |
plain_text | boolean | Whether plain-text mode is enabled |
characters_billed | number | Billing basis for this request (character count of content) |
charged | string | Points consumed by this operation, calculated from the rate. This value reflects usage: usage already covered by an unlimited plan is still reported here |
idempotency_key | string | The idempotency key from the request, echoed verbatim (for reconciliation / correlation) |
billed | boolean | Whether this request was actually billed — false for retries (free) and empty generation results. Reconcile against this field; do not simply sum characters_billed |
summary_language | string | The language actually used for this summary (BCP 47). The language value if one was sent; otherwise zh-TW (the default of language). Also present for free retries and empty generation results, and always has a value (added in v1.16.5) |
prompt_snapshot | string | Present only in custom mode — the customer's prompt exactly as sent |
truncated | boolean | Present only when the summary could not be produced in full (the value is always true). The field is absent entirely when the summary is complete. When present, final_content is not a complete summary: the summary reached the output length limit, or generation reached the processing time limit (only the part completed so far is returned). Both cases are billed as usual. You may want to alert the user, or regenerate with a more concise summary template |
Differences from the Regenerate Summary
doneevent: this endpoint has notask_id(not tied to a recording) and nopersisted(never stored), and adds theidempotency_keyfield.The
characters_billedandchargedfields carry the same meaning as in Regenerate Summary. The difference is that this endpoint always includes them (they can always be determined), whilebilledisfalsefor free retries and empty generation results. Reconcile againstbilled.
Endpoint-Specific Error Codes
Before the stream starts (real HTTP status code + JSON):
| Error code | HTTP | Description | Suggested handling |
|---|---|---|---|
validation_failed | 422 | Parameter validation failed (missing fields, length limits exceeded, mode/field mismatch, etc.) | Fix the fields per data.details.errors |
summary_idempotency_key_conflict | 409 | The same idempotency_key was already used with a different request (content or any summary parameter differs) | Use a new idempotency_key; retries must carry exactly the same fields as the original request |
sse_template_not_found | 404 | Summary template not found (builtin mode; template missing or deactivated) | Verify template |
summary_text_empty | 400 | content has nothing to summarize (all whitespace) | Provide valid content |
stt_quota_exceeded | 402 | Available credits are insufficient for this request's charge | Top up and retry |
too_many_requests | 429 | The free-regeneration limit for this idempotency_key has been reached | Free retries exist for the "already charged but delivery failed" case; they are limited in number and counted over a rolling 24-hour window. To generate again right away, use a new idempotency_key (charged as a new request). Note: This state does not clear within seconds — do not enter a short-interval retry loop on it |
After the stream starts (HTTP 200 + SSE error event):
| Error code | Description | Suggested handling |
|---|---|---|
sse_summary_regeneration_failed | Generation failed (including content filtered by the LLM service, or a stream that stalls or does not end normally; the response does not include internal error details). The summary_regeneration fragments received before it are not a complete result; discard them | Retry later; if the content was filtered, adjust the content or prompt. Failures are not billed and do not claim the identifier |
Version: V1.24.1 Last Updated: 2026-09-28