Streaming
NoPII fully supports streaming responses ("stream": true) for both OpenAI and Anthropic endpoints. Token replacement happens in real time as SSE chunks arrive from the LLM provider.
How streaming works
When streaming is enabled, the LLM provider sends the response as a series of Server-Sent Events (SSE). NoPII processes each chunk in real time:
- 1SSE chunks arrive from the LLM provider
- 2Each chunk is buffered and tokens are replaced with original PII
- 3Detokenized text is verified before being emitted
- 4Detokenized text is emitted as SSE chunks to your application
- 5On stream end, the buffer is flushed to ensure no data is lost
Examples
Python (OpenAI)
from openai import OpenAI
client = OpenAI(base_url="https://api.nopii.co", api_key="sk-your-key")
stream = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Summarize the case for John Smith"}],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Python (Anthropic)
from anthropic import Anthropic
client = Anthropic(base_url="https://api.nopii.co", api_key="sk-ant-your-key")
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Draft a letter for Jane Doe at 123 Main St"}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)cURL
curl -N https://api.nopii.co/chat/completions \
-H "Authorization: Bearer sk-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"stream": true,
"messages": [{"role": "user", "content": "Summarize for John Smith SSN 123-45-6789"}]
}'The response is a standard SSE stream. PII in the LLM output is detokenized in each chunk before it reaches your application.
How buffering works
Tokens can be split across SSE chunk boundaries. NoPII buffers a small number of characters to ensure tokens like [PERSON: aBcDeFgH12] are correctly detected even when the closing bracket arrives in a separate chunk.
This buffering adds no meaningful latency - it only affects the last few characters of each chunk, and all remaining content is flushed when the stream ends.
Response format
The streaming response format is identical to the upstream provider's SSE format. Your existing streaming parsing code works without changes - NoPII only modifies the text content within each chunk.
| Provider | Stream end signal |
|---|---|
| OpenAI | data: [DONE] |
| Anthropic | event: message_stop |
Usage metering
Streaming requests are fully metered for billing, just like non-streaming requests. NoPII captures token usage from the final chunk of the stream (OpenAI) or from message_start and message_delta events (Anthropic). See Billing for how tokens are counted.
Related
- API Reference - Full endpoint documentation including streaming parameters
- Error Handling - What happens when errors occur during a stream