Streaming
Receive partial output as it is generated.
Streaming allows an application to receive content while the model is generating it. It is useful for chat interfaces and longer responses. Before enabling it, confirm that the selected model and interface support streaming.
Enable streaming
Set the following field in a chat request:
{
"stream": true
}Complete request example:
curl --no-buffer --silent --show-error \
"https://api.vergora.ai/v1/chat/completions" \
-H "Authorization: Bearer ${VERGORA_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_STREAMING_MODEL_ID",
"messages": [
{
"role": "user",
"content": "Explain streaming in three sentences."
}
],
"stream": true
}'Read events
A compatible SSE response may contain events such as:
data: {"choices":[{"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}]}
data: [DONE]Use the actual interface response as the source of truth for event structure and completion signals.
Process the stream correctly
- Read data by SSE event boundaries.
- Do not treat one network chunk as a complete JSON value.
- Combine text deltas in order.
- Handle empty deltas, completion events, and in-stream errors.
- Mark the answer as complete only after receiving the expected completion signal.
Interruptions, retries, and cost
After a network interruption or client timeout, mark received content as incomplete. Sending the request again normally creates another request and may create another charge. Check the original request in Requests before retrying.
The absence of usage fields in the stream does not mean that no usage occurred. Use the request record for final usage and cost.