Join our community of builders on Discord!

Streaming

Set stream: true and the answer arrives as OpenAI server-sent events while the worker generates it. OpenAI SDKs and streaming UIs built for them work unchanged.
CodeTYPESCRIPT

The events

The response is text/event-stream. Each event is data: <JSON>:
  1. The first chunk's delta carries role: "assistant" and the first text.
  2. Each later chunk's delta carries more content.
  3. The last chunk has an empty delta, finish_reason: "stop", and the lightchain object naming the on-chain job (see Verifying an Answer).
  4. data: [DONE].
A worker that does not stream sends its whole answer at once: then a single chunk carries the role, the text, finish_reason and lightchain together. The x-lightchain response header carries the same lightchain object from the start, for clients that only read headers.
CodeTEXT

Waiting for the first chunk

Nothing is sent until the first text exists. That can take a while: when none of the key's sessions for the model is free, the call waits for a worker to take on a new one, and a reasoning model thinks before it answers (its reasoning is not forwarded). If a proxy between you and the API cuts idle connections, give it a timeout of a few minutes.

Errors

  • Before the first chunk, a failure is an ordinary HTTP error with its status and body, exactly as without streaming: a 402 still carries accepts, a 429 still carries retry-after.
  • After the first chunk, the status is already 200, so the stream ends with an error event instead of [DONE]:
    CodeTEXT
    OpenAI SDKs raise it as an error while you iterate. Send the request again.
stream_diverged is a rare error code: the worker restarted its answer after the first chunks went out, so the text you received is not the answer it committed on chain. Discard it and send the request again.

Streamed text and the committed answer

The chunks are a best-effort live view. The answer of record is the one the worker commits on chain; the API checks what it streamed against it, sends any part the live view dropped before the last chunk, and reports stream_diverged if they disagree. Closing the connection does not cancel the job: it runs to the end and is paid for.