Preserve Anthropic thinking signatures by routing thinking-capable /messages requests through Responses API - #1
Conversation
Co-authored-by: Godzilla675 <131464726+Godzilla675@users.noreply.github.com>
Co-authored-by: Godzilla675 <131464726+Godzilla675@users.noreply.github.com>
/messages requests through Responses API
There was a problem hiding this comment.
Pull request overview
Routes Anthropic /messages requests that involve thinking state through the Copilot Responses API so encrypted reasoning signatures can be round-tripped (including streaming), preserving continuity across turns.
Changes:
- Added Anthropic ↔ Responses translation for signed thinking blocks via Responses
reasoningitems (includinginclude: ["reasoning.encrypted_content"]). - Added Responses streaming event translation into Anthropic SSE thinking/signature deltas.
- Updated
/messageshandler to conditionally use Responses API, and extended Responses typings + added tests.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/messages-responses-translation.test.ts | Adds coverage for request/response and streaming translation of reasoning/thinking signatures. |
| src/services/copilot/create-responses.ts | Extends Responses payload/types to include reasoning input/output and include. |
| src/routes/messages/responses-translation.ts | Implements non-streaming Anthropic ↔ Responses translation, including signature parsing/building. |
| src/routes/messages/responses-stream-translation.ts | Implements Responses stream → Anthropic SSE translation for thinking/text/tools. |
| src/routes/messages/handler.ts | Routes thinking-capable /messages requests through Responses; wires translators + streaming. |
| src/routes/messages/anthropic-types.ts | Extends Anthropic thinking block type to optionally include a signature. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| function handleToolArgumentsDelta( | ||
| parsedEvent: { | ||
| output_index?: unknown | ||
| delta?: unknown | ||
| }, | ||
| state: ResponsesStreamState, | ||
| ): Array<AnthropicStreamEventData> { | ||
| if ( | ||
| typeof parsedEvent.output_index !== "number" | ||
| || typeof parsedEvent.delta !== "string" | ||
| ) { | ||
| return [] | ||
| } | ||
|
|
||
| const blockIndex = state.blockIndexByKey.get( | ||
| `tool:${parsedEvent.output_index}`, | ||
| ) | ||
| if (blockIndex === undefined) { | ||
| return [] | ||
| } | ||
|
|
||
| return [ | ||
| { | ||
| type: "content_block_delta", | ||
| index: blockIndex, | ||
| delta: { | ||
| type: "input_json_delta", | ||
| partial_json: parsedEvent.delta, | ||
| }, | ||
| }, | ||
| ] | ||
| } | ||
|
|
||
| function handleToolArgumentsDone( | ||
| parsedEvent: { | ||
| output_index?: unknown | ||
| arguments?: unknown | ||
| }, | ||
| state: ResponsesStreamState, | ||
| ): Array<AnthropicStreamEventData> { | ||
| if ( | ||
| typeof parsedEvent.output_index !== "number" | ||
| || typeof parsedEvent.arguments !== "string" | ||
| ) { | ||
| return [] | ||
| } | ||
|
|
||
| const blockIndex = state.blockIndexByKey.get( | ||
| `tool:${parsedEvent.output_index}`, | ||
| ) | ||
| if (blockIndex === undefined || state.blockHasDelta.has(blockIndex)) { | ||
| return [] | ||
| } | ||
|
|
||
| state.blockHasDelta.add(blockIndex) | ||
| return [ | ||
| { | ||
| type: "content_block_delta", | ||
| index: blockIndex, | ||
| delta: { | ||
| type: "input_json_delta", | ||
| partial_json: parsedEvent.arguments, | ||
| }, | ||
| }, | ||
| ] | ||
| } |
There was a problem hiding this comment.
handleToolArgumentsDelta emits input_json_delta events but never marks the tool block as having received deltas. As a result, when response.function_call_arguments.done arrives, handleToolArgumentsDone will also emit the full arguments again, duplicating tool input in the Anthropic stream. Track deltas for tool blocks (e.g., add the block index to blockHasDelta in the delta handler, or use a dedicated set for tool-arg deltas) so .done only backfills when no deltas were seen.
| function openBlock(params: { | ||
| state: ResponsesStreamState | ||
| key: string | ||
| contentBlock: | ||
| | { type: "thinking"; thinking: string } | ||
| | { type: "text"; text: string } | ||
| | { | ||
| type: "tool_use" | ||
| id: string | ||
| name: string | ||
| input: Record<string, unknown> | ||
| } | ||
| events: Array<AnthropicStreamEventData> | ||
| }): number { | ||
| const { state, key, contentBlock, events } = params | ||
| let blockIndex = state.blockIndexByKey.get(key) | ||
| if (blockIndex === undefined) { | ||
| blockIndex = state.nextContentBlockIndex | ||
| state.nextContentBlockIndex += 1 | ||
| state.blockIndexByKey.set(key, blockIndex) | ||
| } | ||
|
|
||
| if (state.openBlockKey === key) { | ||
| return blockIndex | ||
| } | ||
|
|
||
| closeOpenBlock(state, events) | ||
| events.push({ | ||
| type: "content_block_start", | ||
| index: blockIndex, | ||
| content_block: contentBlock, | ||
| }) | ||
| state.openBlockKey = key | ||
| state.openBlockIndex = blockIndex | ||
| return blockIndex | ||
| } |
There was a problem hiding this comment.
openBlock reuses the same index for a given key even after a content_block_stop, which means interleaved Responses events can cause multiple content_block_start events for the same index (e.g., switching from thinking:0 to text:1:0 and back). The existing chat-completions stream translator only ever increments indices and never restarts an earlier index; emitting multiple starts for the same index is likely incompatible with Anthropic’s streaming protocol and can break clients. Consider buffering/reordering interleaved deltas or allocating a new monotonically increasing index each time a block is (re)started.
| return { | ||
| model: payload.model, | ||
| input, | ||
| stream: payload.stream, |
There was a problem hiding this comment.
translateAnthropicToResponses forwards payload.model directly, but the existing /messages → chat-completions translation normalizes certain Claude subagent model variants because Copilot doesn’t support them (see translateModelName in non-stream-translation.ts). Without similar normalization here, the same request may succeed on the chat-completions path but fail when routed to Responses. Reuse the same model normalization logic (or centralize it) before calling the Responses API.
| const responsesPayload = translateAnthropicToResponses(anthropicPayload) | ||
| consola.debug( | ||
| "Translated Responses payload:", | ||
| JSON.stringify(responsesPayload), | ||
| ) | ||
|
|
||
| const response = await createResponses(responsesPayload) | ||
|
|
||
| if (isNonStreamingResponse(response)) { | ||
| consola.debug( | ||
| "Non-streaming response from Responses API:", | ||
| JSON.stringify(response).slice(-400), | ||
| ) | ||
| const anthropicResponse = translateResponsesToAnthropic(response) | ||
| consola.debug( | ||
| "Translated Anthropic response:", | ||
| JSON.stringify(anthropicResponse), | ||
| ) | ||
| return c.json(anthropicResponse) | ||
| } | ||
|
|
||
| consola.debug("Streaming response from Responses API") | ||
| return streamSSE(c, async (stream) => { | ||
| const streamState = createResponsesStreamState() | ||
|
|
||
| for await (const rawEvent of response) { | ||
| consola.debug("Responses raw stream event:", JSON.stringify(rawEvent)) | ||
| if (!rawEvent.data || rawEvent.data === "[DONE]") { | ||
| continue | ||
| } | ||
|
|
||
| const events = translateResponsesStreamEvent( | ||
| JSON.parse(rawEvent.data) as { | ||
| type?: string | ||
| [key: string]: unknown | ||
| }, | ||
| streamState, | ||
| ) | ||
|
|
||
| for (const event of events) { | ||
| consola.debug("Translated Anthropic event:", JSON.stringify(event)) | ||
| await stream.writeSSE({ | ||
| event: event.type, | ||
| data: JSON.stringify(event), | ||
| }) |
There was a problem hiding this comment.
The new debug logs in the Responses path print full request/stream payloads, including reasoning.encrypted_content / thinking signatures and potentially user content or base64 images. These values are effectively replayable state and can be sensitive; logging entire events can also be very noisy in production. Consider redacting encrypted_content/signature fields and truncating payloads (similar to existing .slice(-400) usage) or gating these logs behind an explicit verbose/debug flag.
|
@copilot address code review |
Co-authored-by: Godzilla675 <131464726+Godzilla675@users.noreply.github.com>
Co-authored-by: Godzilla675 <131464726+Godzilla675@users.noreply.github.com>
Addressed in eb06825 and a1edfd5. I fixed the Responses stream review issues (tool-arg delta duplication, reopened block index reuse, Claude model normalization on the Responses path, and sensitive Responses-path debug logging), added regression coverage, and resolved the shared Responses type issue that full lint exposed afterward. Warning Firewall rules blocked me from connecting to one or more addresses (expand for details)I tried to connect to the following addresses, but was blocked by firewall rules:
If you need me to access, download, or install something from one of these locations, you can either:
|
⚡ Quickly spin up Copilot coding agent tasks from anywhere on your macOS or Windows machine with Raycast.