Skip to main content
GET /threads/{threadId}/stream and GET /tasks/{taskId}/stream hold a connection open and push what GET /threads/{threadId}/messages would page: the same Message entries, in the same order, as their rows commit. On top of that they carry the turn in progress, so you can show the agent’s reply as it is written and the tools it is calling, without polling.
The response is text/event-stream. Each frame is an event: name and one JSON data: line; transcript frames also carry an id: line, the cursor you resume from. A retry: 2000 line follows the first frame, the delay to wait before reconnecting after a dropped connection.
The browser EventSource cannot send an Authorization header, so read the stream with fetch (or any HTTP client) and parse the frames yourself. Ignore event names you do not know: the vocabulary grows over time and a newer server may send names this page does not list.

Events

Live frames show what the transcript will contain, before it commits. Two rules connect them:
  • A transcript frame with source: "assistant" carries the callId of the delivery call it came from. When it arrives, replace every assistant fragment you accumulated under that callId with the entry’s text.
  • A transcript frame with source: "tool" has the same id as the tool_call frame’s callId. Replace the live frame with the entry.
Fragments whose callId never settles into a transcript frame belong to a turn that was interrupted or a reply whose delivery failed; discard them at the next run end. A turn end with outcome dropped means the server discarded live frames of that turn because you were reading too slowly or the turn was too large to replay; the transcript frames still arrive in full. Treat live frames as a preview and transcript frames as the record. The stream carries nothing the transcript hides: tool activity is the rendered summary, never arguments or results, and the only text that streams is the reply being delivered to you, never the model’s internal reasoning.

Resuming

Store the last id: you read. On reconnect, send it as the Last-Event-ID header (or the after query parameter, when you cannot set headers) and the stream replays every transcript entry after that cursor, then the turn in progress, then continues live. Cursors are the transcript’s event ids and never expire.
A value that is not a valid cursor answers 400 capy/InvalidRequest before the stream starts.

Waiting for one reply

Pass until=run to close the stream once the run in progress ends. You get every frame of the run, the transcript entries it committed, then done. When the thread is not working and has nothing queued, done arrives right away.
This replaces the poll-until-idle loop: send a message, open the stream with until=run, and read until it closes. The reply is the last transcript entry with source: "assistant". Without a cursor the stream first replays the whole transcript; pass the last id: you read as Last-Event-ID or after to skip it.

Task streams

GET /tasks/{taskId}/stream carries one task’s own turns with the same frames. Its status frame has the task shape, and it closes with done when the task’s status reaches done. A thread stream does not carry its tasks’ tokens, only their lifecycle as task frames; open a task’s stream when you want to follow its work.

Limits

  • The stream authenticates like every other request and answers 404 capy/ThreadNotFound or capy/TaskNotFound where the JSON reads do. Every error before the first byte is an ordinary JSON error; after that, failures arrive as an error frame.
  • Your API key is re-verified every 15 seconds. A key revoked, expired, or disabled while a stream is open closes it with error code unauthorized within that interval.
  • Each API server holds a fixed number of open streams. Above it a new connection answers 429 capy/RateLimited with retryAfterSeconds; wait and retry.
  • A connection can be cut without a closing frame when a server restarts. Treat any close without done or error as a resume: wait retry milliseconds and reconnect with Last-Event-ID.
  • The heartbeat keeps idle connections alive through proxies. If your client has its own idle timeout, set it above 15 seconds.