1. The prompt
Create a prompt (sayassistant) with read-write memory — it sees the conversation
so far and appends its replies to it (see
Memory & Threads). Give it your system instructions and
an input variable for the user’s message, then publish it to production.
2. The backend
Two endpoints, both thin. One mints the stream credential, one triggers runs. Your API key stays here; the browser never sees it.randomUUID()), keep it in your session or database, and pass the same one to both
endpoints. The thread doesn’t need to exist before the first run; minting a stream token
for it first is exactly the right order.
3. The frontend
Connect before sending the first message — a fresh subscription starts at the live tip, so open the stream first and you see every token from the beginning.send triggers a run, the run streams into runs[runId].text,
and status flips to done when it finishes. Reconnects, retries, and workflow
segments are handled inside the hook.
When a terminal state carries
gapped: true, the streamed text may be incomplete —
fetch the run through your backend (pj.getPromptRun(runId)) and swap in the
authoritative output. With connect-before-send this is rare, but handling it is what
makes the UI trustworthy.Showing the user’s side
The stream carries the model’s output; your user’s messages you already have locally. Render them optimistically from your own state, keyed however you like — the runs map and your local messages interleave naturally by time. On a full page reload, rebuild the transcript from your backend (the thread’s runs via the API, or your own message store) and let the stream take over for new messages.Workflows instead of a single prompt
SwaprunPrompt for runWorkflow and nothing else changes — every prompt node streams
into the same connection. If the workflow has internal lanes (research, tool use,
synthesis), put the user-facing node on a dedicated
channel and pass channels: ['support'] to the
hook so only that lane renders.