# LLM Streaming Stops Early: Diagnose SSE and Completion

> Debug interrupted LLM streams without treating partial text as success. Separate network chunks, protocol events, finish reasons and final usage records.

_Published 2026-09-29 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![End-to-end AI request latency trace split into client, gateway, provider, retries, and generation](https://keeprouter.com/editorial/blog/ai-gateway-latency-guide.png)

_Distinguish the first visible event from completion of the entire task. Conceptual illustration._

An LLM stream is complete only when the protocol reaches a valid terminal state for the operation. Receiving some text or an HTTP 200 is not enough. A connection can close after partial output, a token limit can truncate the answer, or a parser can discard data split across network chunks.

This guide is for an application that already displays streaming text but occasionally produces cut-off answers or unexplained empty results. Start by identifying the earliest boundary where evidence disappears, rather than assuming the model stopped thinking.

## Separate transport chunks from protocol events

[MDN's Server-Sent Events guide](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events) describes the event framing. A network read is not guaranteed to contain one complete event. It may contain half an event or several events, and UTF-8 text can also span reads.

Use a streaming decoder and a protocol-aware parser, or the official SDK's streaming interface. Do not call `JSON.parse` on each raw network chunk. Keep incomplete data buffered until a full event is available. A parser that works on localhost can fail under different packet boundaries without any change to the model.

| Layer | What it can tell you | What it cannot prove |
| --- | --- | --- |
| HTTP status | Whether the initial response was accepted | The generation finished |
| Network read | Bytes arrived | A full SSE event arrived |
| Text delta | Partial content exists | The answer is complete |
| Finish event | A protocol-specific ending occurred | Every ending is a successful answer |
| Usage record | Measured fields were provided | Missing fields should be zero |

Keep these states explicit in application code. Do not use the presence of non-empty text as the only success condition.

## Know the terminal events for your API

Chat Completions, Responses and Anthropic Messages do not share one universal event format. [Claude's streaming documentation](https://platform.claude.com/docs/en/build-with-claude/streaming) describes message and content-block events, including error handling. An application that expects a Chat Completions sentinel cannot safely consume a different event stream unchanged.

For the selected protocol, identify the final event, finish reason and where usage appears. Test normal completion, output-limit termination, refusal, tool handoff and explicit error separately. A tool handoff may be a normal end to one model request while the larger user task remains unfinished.

The following is a local state-machine fixture, not an SSE parser or a real provider response:

```python
def result_state(text, finish_reason, transport_error=False):
    if transport_error:
        return "interrupted"
    if finish_reason == "length":
        return "truncated"
    if finish_reason == "tool_calls":
        return "needs_tool"
    if finish_reason == "stop":
        return "complete"
    return "incomplete"

assert result_state("Half an answer", None) == "incomplete"
assert result_state("A full answer", "stop") == "complete"
assert result_state("Partial JSON", "length") == "truncated"
```

Adapt the states to the exact API contract. Keep empty normal completions distinguishable from transport errors, and decide whether an empty answer satisfies your product's task. A terminal state alone cannot grade the answer's usefulness.

## Reproduce the failure at one boundary

First send the same short fixture without streaming. If that succeeds, try streaming through the backend without rendering it. Then enable the frontend renderer. This narrows the investigation to model output, backend parsing, proxy behavior or UI state.

Use a prompt with a modest requested answer length. Log event types and timestamps without logging private text. Capture when the request began, the first event arrived, the first visible text appeared and the terminal event arrived. Those measurements distinguish a slow start from a mid-stream pause.

When testing a reverse proxy, check its buffering and timeout behavior. An idle timeout and a total request deadline are different controls. Raising one will not fix a parser that loses partial frames. Likewise, disabling timeouts entirely can leave abandoned requests consuming resources.

## Handle cancellation honestly

When a user presses Stop, abort the request through the relevant client mechanism and show the answer as cancelled. Do not relabel it as a complete answer because it contains useful text. Preserve the partial draft only if that suits the product, and make its state visible.

Cancellation does not prove that upstream computation and billing stopped at the exact same instant. Reconcile final usage when available. If the stream ended before a final usage record, mark usage unknown until a reliable record is available; do not manufacture a zero-cost result.

For tool calls, never execute arguments collected from an incomplete stream. The [tool-loop guide](/blog/llm-tool-calling-loop) explains why incomplete arguments and duplicate actions require separate safeguards.

## Retry without corrupting the conversation

A retry after partial text can create two overlapping answers. Keep the attempts separate and let the application choose which result is active. For structured output, do not concatenate a second response onto a truncated JSON object and hope it becomes valid.

If the request included a side-effecting tool, check whether the operation already completed before replaying the task. A model retry and a business-action retry are not the same thing.

Use the [error reference](/docs/errors) to classify returned errors and the [API documentation](/api/docs) for the selected route. Before changing models, reproduce the small fixture through the exact production path. Then compare completion rate, accepted answers and total charges, using the [latency guide](/blog/ai-gateway-latency-guide) to keep first-token timing separate from full task completion.

## Frequently asked questions

### Does HTTP 200 mean the streamed answer finished?

No. It describes the initial response. Track the protocol's terminal event and finish reason, and distinguish truncation, tool handoff and cancellation.

### Can I parse each network chunk as JSON?

No. Network chunk boundaries do not match SSE event boundaries. Buffer and decode complete protocol events or use an appropriate SDK parser.

### Is a cancelled stream free?

Do not assume that. Some computation may already have occurred. Use the destination's billing rules and reliable usage records rather than the length of visible text.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [MDN Server-Sent Events](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events)
2. [Claude streaming events](https://platform.claude.com/docs/en/build-with-claude/streaming)

## Related guides

- [LLM Tool Calling: Build the Full Request-Result Loop](https://keeprouter.com/blog/llm-tool-calling-loop.md)
- [OpenAI-Compatible API 429 Errors: Limits, Quota and Retry](https://keeprouter.com/blog/openai-compatible-api-429-errors.md)
- [AI gateway latency guide: measure the hop, the provider and the rescued request](https://keeprouter.com/blog/ai-gateway-latency-guide.md)

## Check the behavior your application needs

Use the documented request format and controls when applying the example to your application.

[Read the product reference](https://keeprouter.com/docs/errors)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
