SSE API

Ad-hoc Summary SSE

Connection Info

ItemValue
Base pathhttps://vas-poc.vurbo.ai/api/v1/sse
ProtocolHTTP + Server-Sent Events (SSE)
Data formattext/event-stream
AuthenticationHeader X-API-Key: {KEY} (this endpoint does not accept the key via query string)

Note: This endpoint only accepts POST (JSON body), so the browser-native EventSource API cannot be used. Call it with the fetch API plus ReadableStream, or any HTTP client that can read streamed responses.


Endpoint Overview

MethodEndpointPersists ResultBilledPurpose
POST/api/v1/sse/summaryNoYesGenerate a summary from arbitrary text supplied in the request (not tied to any recording)

Difference from Regenerate Summary: regeneration always uses the transcript stored on the server as input, while this endpoint uses the content supplied in the request — for content the server does not have, such as the full transcript of several merged recordings, or a transcript edited by the user. The result is never stored; it is only streamed back to the client.

Added in v1.9.1. An earlier endpoint once existed at the same path (removed in V1.8.0); this endpoint is a brand-new contract — authentication, billing, and duplicate-request rules all differ, so do not reuse old integration code.


Request Parameters (JSON body)

ParameterTypeRequiredConstraintsDescription
contentstringYes≤200,000 charactersThe full text to summarize
idempotency_keystringYes≤64 characters, A-Z a-z 0-9 . _ -Duplicate-request identifier that prevents double billing (see "Duplicate Requests and Billing Guarantees" below)
modestringYes"builtin" | "custom"Summary mode: built-in template or custom prompt
templatestringRequired for builtin / forbidden for customA valid built-in template slugBuilt-in template slug
promptstringRequired for custom / forbidden for builtin≤3000 charactersThe customer's full prompt
promptSlugstringRequired for custom / forbidden for builtin≤64 characters, Unicode, no control charactersCustomer-defined identifier (returned as-is, not processed)
languagestringNo-Output language code for the summary (e.g. zh-TW, en-US); defaults to zh-TW when omitted
plainTextbooleanNoDefaults to falseRequest plain-text output (Markdown formatting symbols are removed automatically)

Mutual-exclusion rules (same as Regenerate Summary):

  • With mode=builtin, prompt and promptSlug must not be sent
  • With mode=custom, template must not be sent, while prompt and promptSlug are required

Error Response Model (differs from other SSE endpoints — please note)

This endpoint is intended for backend integrations. Errors before the stream starts always return a real HTTP status code with a JSON body, instead of the "HTTP 200 + error event" convention used by other SSE endpoints:

PhaseResponse form
Authentication failure (401/403), validation failure (422), template missing or deactivated (404), all-whitespace content (400), insufficient credit (402), duplicate-request identifier conflict (409), rate limit exceeded (429)Real HTTP status code + JSON error body
After the stream starts (generation failure, filtered content, etc.)HTTP 200 + SSE error event

Authentication only accepts the X-API-Key header; API keys in the query string are not accepted.

The 422 JSON body carries per-field error messages under data.details.errors, so your program can act on them:

{
  "type": "error",
  "data": {
    "error_code": "validation_failed",
    "message": "Validation failed",
    "details": { "errors": { "content": ["The content parameter is required"] } }
  }
}

Billing

  • Billed by the character count of content: 0.1 credits per 1,000 characters (rounded up below 1,000; same rate as Meeting Summary and Regenerate Summary).
  • Billed only on successful generation; failed generation (including filtered content, or a stream that stalls or does not end normally) is not billed. A summary that is incomplete, with truncated in done, still counts as successful and is billed as usual.
  • Example: a 35,000-character content = 3.5 credits; 1 character = 0.1 credits (minimum billing unit).

Duplicate Requests and Billing Guarantees

idempotency_key is supplied by the client (we recommend a stable identifier from your system, such as a merge-batch ID or revision ID). The system compares the entire request — content plus every summary parameter (mode / template / prompt / promptSlug / language / plainText) — to decide whether two calls are retries of the same piece of work:

ScenarioBehavior
Retry with the same identifier + an identical requestNot billed again, but the summary is regenerated (the result may differ from last time; the previous result is not replayed); retries still work even when the account balance has run out (the original request was already billed)
Same identifier + any field differs (content or any summary parameter)Returns 409 summary_idempotency_key_conflict; nothing is generated or billed
Retry after a failed first requestFailure does not claim the identifier — retrying (even with different content) is treated as a brand-new request and billed normally

Retries must carry exactly the same fields as the original request; when summarizing new content or switching prompts, always send a new idempotency_key.

idempotency_key is only valid within a single API key: the same identifier on different API keys is unrelated and billed separately. Re-sending the same piece of work through a backup key counts as two requests.

Edge case: if the same identifier is used concurrently for two different requests (a client error), the first to complete claims it; the other is still billed and delivered normally, but any later retry with that identifier (regardless of content) returns 409 — switch to a new identifier.


Request Examples

builtin mode

curl -N -X POST "https://vas-poc.vurbo.ai/api/v1/sse/summary" \
  -H "X-API-Key: vas_aB3dE5fG7hI9jK1lM3nO5pQ7rS9tU1vW" \
  -H "Content-Type: application/json" \
  -d '{
    "content": "(full transcript merged from multiple recordings…)",
    "idempotency_key": "merge-batch-20260813-001",
    "mode": "builtin",
    "template": "meeting",
    "language": "en-US",
    "plainText": true
  }'

custom mode

curl -N -X POST "https://vas-poc.vurbo.ai/api/v1/sse/summary" \
  -H "X-API-Key: vas_..." \
  -H "Content-Type: application/json" \
  -d '{
    "content": "(full transcript after user edits…)",
    "idempotency_key": "revision-8f31c2",
    "mode": "custom",
    "prompt": "Extract the decisions and action items from the transcript…",
    "promptSlug": "acme-meeting-v2",
    "language": "en-US"
  }'

Event Sequence

1. connected              → connection confirmed
2. summary_regeneration   → summary fragments (repeated N times, cumulative)
3. done                   → generation complete
   or
3. error                  → generation failed (sse_summary_regeneration_failed); no done is sent and nothing is billed

connected

{ "message": "Ad-hoc summary stream connected (mode: builtin)" }

summary_regeneration

{ "text": "The meeting covered the following topics:\n1. Product development progress", "is_final": false }
FieldTypeDescription
textstringCumulative summary content (with plainText=true, the is_final=true text is the cleaned plain text)
is_finalbooleanWhether this is the final result

done

{
  "tokens_used": 123,
  "final_content": "The meeting... (full summary content)",
  "mode": "custom",
  "template": "acme-meeting-v2",
  "plain_text": true,
  "characters_billed": 35000,
  "charged": "3.5",
  "idempotency_key": "revision-8f31c2",
  "billed": true,
  "summary_language": "zh-TW",
  "prompt_snapshot": "Extract the decisions and action items from the transcript…"
}
FieldTypeDescription
tokens_usednumberTotal token usage
final_contentstringFull summary content (cleaned plain text when plainText=true)
modestringSummary mode: "builtin" or "custom"
templatestringThe identifier actually used — builtin → built-in template slug; custom → customer-defined identifier
plain_textbooleanWhether plain-text mode is enabled
characters_billednumberBilling basis for this request (character count of content)
chargedstringPoints consumed by this operation, calculated from the rate. This value reflects usage: usage already covered by an unlimited plan is still reported here
idempotency_keystringThe idempotency key from the request, echoed verbatim (for reconciliation / correlation)
billedbooleanWhether this request was actually billed — false for retries (free) and empty generation results. Reconcile against this field; do not simply sum characters_billed
summary_languagestringThe language actually used for this summary (BCP 47). The language value if one was sent; otherwise zh-TW (the default of language). Also present for free retries and empty generation results, and always has a value (added in v1.16.5)
prompt_snapshotstringPresent only in custom mode — the customer's prompt exactly as sent
truncatedbooleanPresent only when the summary could not be produced in full (the value is always true). The field is absent entirely when the summary is complete. When present, final_content is not a complete summary: the summary reached the output length limit, or generation reached the processing time limit (only the part completed so far is returned). Both cases are billed as usual. You may want to alert the user, or regenerate with a more concise summary template

Differences from the Regenerate Summary done event: this endpoint has no task_id (not tied to a recording) and no persisted (never stored), and adds the idempotency_key field.

The characters_billed and charged fields carry the same meaning as in Regenerate Summary. The difference is that this endpoint always includes them (they can always be determined), while billed is false for free retries and empty generation results. Reconcile against billed.


Endpoint-Specific Error Codes

Before the stream starts (real HTTP status code + JSON):

Error codeHTTPDescriptionSuggested handling
validation_failed422Parameter validation failed (missing fields, length limits exceeded, mode/field mismatch, etc.)Fix the fields per data.details.errors
summary_idempotency_key_conflict409The same idempotency_key was already used with a different request (content or any summary parameter differs)Use a new idempotency_key; retries must carry exactly the same fields as the original request
sse_template_not_found404Summary template not found (builtin mode; template missing or deactivated)Verify template
summary_text_empty400content has nothing to summarize (all whitespace)Provide valid content
stt_quota_exceeded402Available credits are insufficient for this request's chargeTop up and retry
too_many_requests429The free-regeneration limit for this idempotency_key has been reachedFree retries exist for the "already charged but delivery failed" case; they are limited in number and counted over a rolling 24-hour window. To generate again right away, use a new idempotency_key (charged as a new request). Note: This state does not clear within seconds — do not enter a short-interval retry loop on it

After the stream starts (HTTP 200 + SSE error event):

Error codeDescriptionSuggested handling
sse_summary_regeneration_failedGeneration failed (including content filtered by the LLM service, or a stream that stalls or does not end normally; the response does not include internal error details). The summary_regeneration fragments received before it are not a complete result; discard themRetry later; if the content was filtered, adjust the content or prompt. Failures are not billed and do not claim the identifier

Version: V1.24.1 Last Updated: 2026-09-28

Copyright © 2026