Streaming

NoPII fully supports streaming responses ("stream": true) for both OpenAI and Anthropic endpoints. Token replacement happens in real time as SSE chunks arrive from the LLM provider.

How streaming works

When streaming is enabled, the LLM provider sends the response as a series of Server-Sent Events (SSE). NoPII processes each chunk in real time:

  1. 1SSE chunks arrive from the LLM provider
  2. 2Each chunk is buffered and tokens are replaced with original PII
  3. 3Detokenized text is verified before being emitted
  4. 4Detokenized text is emitted as SSE chunks to your application
  5. 5On stream end, the buffer is flushed to ensure no data is lost

Examples

Python (OpenAI)

python
from openai import OpenAI

client = OpenAI(base_url="https://api.nopii.co", api_key="sk-your-key")

stream = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Summarize the case for John Smith"}],
    stream=True,
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Python (Anthropic)

python
from anthropic import Anthropic

client = Anthropic(base_url="https://api.nopii.co", api_key="sk-ant-your-key")

with client.messages.stream(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Draft a letter for Jane Doe at 123 Main St"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

cURL

bash
curl -N https://api.nopii.co/chat/completions \
  -H "Authorization: Bearer sk-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "stream": true,
    "messages": [{"role": "user", "content": "Summarize for John Smith SSN 123-45-6789"}]
  }'

The response is a standard SSE stream. PII in the LLM output is detokenized in each chunk before it reaches your application.

How buffering works

Tokens can be split across SSE chunk boundaries. NoPII buffers a small number of characters to ensure tokens like [PERSON: aBcDeFgH12] are correctly detected even when the closing bracket arrives in a separate chunk.

This buffering adds no meaningful latency - it only affects the last few characters of each chunk, and all remaining content is flushed when the stream ends.

Response format

The streaming response format is identical to the upstream provider's SSE format. Your existing streaming parsing code works without changes - NoPII only modifies the text content within each chunk.

ProviderStream end signal
OpenAIdata: [DONE]
Anthropicevent: message_stop

Usage metering

Streaming requests are fully metered for billing, just like non-streaming requests. NoPII captures token usage from the final chunk of the stream (OpenAI) or from message_start and message_delta events (Anthropic). See Billing for how tokens are counted.

Related

  • API Reference - Full endpoint documentation including streaming parameters
  • Error Handling - What happens when errors occur during a stream