Skip to content

fix: improve streaming reliability and max_completion_tokens support - #220

Open
alext-avi wants to merge 1 commit into
ericc-ch:masterfrom
alext-avi:fix/streaming-reliability
Open

alext-avi wants to merge 1 commit into
ericc-ch:masterfrom
alext-avi:fix/streaming-reliability

Conversation

@alext-avi

Copy link
Copy Markdown

Summary

  • Return raw Response from createChatCompletions instead of parsing SSE in the service layer, allowing the OpenAI-compatible handler to pipe streams directly via TransformStream (avoids re-serialization overhead and buffering issues)
  • Anthropic /messages handler now parses SSE at the handler level where event translation is needed
  • Normalize max_completion_tokens to max_tokens for upstream compatibility (some clients send the newer OpenAI field which Copilot doesn't accept)
  • Set Bun idleTimeout: 0 to prevent premature SSE connection kills (default 10s was dropping long-running LLM streams)

Test plan

  • Verified streaming responses work end-to-end with Claude Code via OpenAI-compatible endpoint
  • Verified Anthropic /messages streaming translation still works
  • Verified non-streaming responses are unaffected
  • Confirmed long-running streams no longer get killed by Bun's idle timeout

🤖 Generated with Claude Code

- Return raw Response from createChatCompletions instead of parsing SSE
  in the service layer, allowing handlers to pipe streams directly
- Pipe OpenAI-compatible streams through a monitoring TransformStream
  for TTFT and throughput logging without re-serializing
- Anthropic /messages handler now parses SSE at the handler level where
  it needs to translate events
- Normalize max_completion_tokens to max_tokens for upstream compat
- Set Bun idleTimeout to 0 to prevent premature SSE connection kills
  (default 10s was dropping long-running LLM streams)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
laizhengbao added a commit to laizhengbao/copilot-api that referenced this pull request Sep 27, 2026
Two fixes in the chat completions handler:

1. Clients that send the newer `max_completion_tokens` parameter
	 (instead of legacy `max_tokens`) receive a 400 from the Copilot API:
	 `max_tokens and max_completion_tokens cannot both be set`.
	 The handler injects its own `max_tokens` whenever the legacy field is
	 unset, and upstream rejects requests carrying both parameters.
	 Injection is now skipped when either parameter is present, and the
	 payload type gains an optional `max_completion_tokens` field so the
	 value is forwarded upstream as intended.

	 Related: ericc-ch#220 addresses the same interop by normalizing
	 `max_completion_tokens` into `max_tokens` instead; this commit takes
	 the passthrough approach. Happy to drop this half if ericc-ch#220 lands first.

2. GitHub's /models endpoint lists some entries WITHOUT a
	 `capabilities` object (observed: `gpt-41-copilot`,
	 `trajectory-compaction`). Requesting one crashed the proxy with a
	 500: `undefined is not an object (evaluating
	 'selectedModel?.capabilities.limits.max_output_tokens')`.
	 `Model.capabilities` is now optional to mirror the real API, and both
	 the handler and the tokenizer lookup guard against its absence.

Tested against a live Copilot endpoint (individual plan):
- `max_completion_tokens`-only request -> 200 (previously 500)
- capability-less model without max_tokens -> clean upstream error, no crash
- plain request without either parameter -> `max_tokens` still injected
- streaming requests unaffected

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant