Long streamed answers now arrive in full
A stream is no longer capped by total response time: it stays alive as long as the model keeps sending and is cut only after a pause longer than 4 minutes. Long answers from reasoning models and agents arrive complete. Dropped streams now show up on the status page.
What changed
A response used to have 2 minutes in total — whether or not the model was still writing. The limit now works differently: what counts is not how long the answer takes, but how long the transfer stays silent.
A stream is cut only if no byte arrives from the model for 4 minutes in a row. Every new part of the answer resets the countdown.
The overall cap for a single response is 30 minutes.
If the model does not start responding at all, we wait 2 minutes and return
504 upstream_timeout.
Who will notice
Reasoning models. Thinking before the first word no longer hits the old limit.
Long generations: large blocks of code, translations, analysis of big files.
Agents and IDE clients — Claude Code, Codex CLI, Cline, Roo Code: multi-step answers with tool calls now run to completion.
What to check on your side
Turn streaming on for long answers — without it, the response still has to arrive within 2 minutes:
{
"model": "anthropic/claude-sonnet-5",
"stream": true,
"messages": [{"role": "user", "content": "..."}]
}Second, your HTTP client timeout. The official OpenAI SDKs wait 10 minutes, but custom wrappers often set 60 seconds: in that case the request is cut on your side, not ours.
The status page is more accurate
A dropped response now counts towards the model's availability on the Status page. A model whose streams break no longer looks perfectly healthy — the page shows which model is best avoided right now.
Billing is unchanged
The answer never started — the request is free.
The stream broke mid-way — you are charged only for what was actually delivered.
All error codes and what to do about each are in the Errors section of the documentation.