Skip to content

feat: serve the Anthropic Messages API on /v1/messages - #372

Merged
Priyanshu-u07 merged 2 commits into
mainfrom
feat/anthropic-messages-surface
Oct 3, 2026
Merged

Priyanshu-u07 merged 2 commits into
mainfrom
feat/anthropic-messages-surface

Conversation

@Priyanshu-u07

Copy link
Copy Markdown
Collaborator

Anything written against the Anthropic Messages API, the Claude SDK and the agent frameworks built on it, could not point at the gateway at all, because only OpenAI-shaped routes existed.

/v1/messages translates the request into the OpenAI shape used internally,runs the ordinary completion path, and translates the answer back. Nothing in that path is aware of the surface, so API key resolution, quotas, rate limiting, routing and logging are the same code rather than a second copy. The diff on existing files is 39 lines: a route, an export and an alias.

A surface is the mirror of a provider adapter. One translates between a client's shape and ours, the other between an upstream's shape and ours:

client  -> [surface in ] -> internal -> [provider out] -> upstream
client <-  [surface out] <- internal <- [provider in ] <- upstream

#368 built the right-hand side. This is the left.

Three things worth a look in review

x-api-key falls back to the bearer header, because Anthropic clients send the key that way and would otherwise fail to authenticate. Sandbox is deliberately excluded from that fallback: it carries a JWT that extract_api_key verifies, and x-api-key must not be a way around it.

The stream translator closes its envelope in a finally. Anthropic wraps text in an explicit envelope that OpenAI carries implicitly, so most of it is synthesised rather than translated — and if upstream dies mid-stream a client is otherwise left waiting on message_stop forever.

Translator state lives on the instance here, unlike #368's, because a surface translator is built per request where provider adapters are cached and shared. Covered by a concurrent-streams test.

Known gaps

Only text content is supported. Image and tool blocks are rejected with a 400 naming the type, rather than dropped, so a request never quietly loses content and answers anyway.

tools is not translated. Largely moot today, since vLLM deployments launch without --enable-auto-tool-choice and so cannot emit structured tool calls either way.

/v1/messages/count_tokens is not routed. vLLM serves it; the SDK onlycalls it when asked to.

Streamed usage.output_tokens counts deltas, not tokens. Upstream reports an exact count only when stream_options.include_usage is requested, which the surface does not control. One delta is one token on vLLM and an
approximation anywhere that batches.

Why this is a draft

Everything upstream of the route is mocked. The translation is verified by the real Anthropic SDK, but no request in these tests reaches vLLM. Draft until it has been run against a live deployment.

One specific thing to check there: cancel a stream mid-response. The tests cover upstream dying, but a client disconnecting is different, Starlette cancels the generator rather than raising into it, and whether the finally still closes the envelope under that cancellation is the one claim here no mock can settle.

Closes #369

Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
@Priyanshu-u07

Copy link
Copy Markdown
Collaborator Author

Verified on a live deployment, g5.xlarge (A10G), kind cluster via deploy/gpu-kind.sh, vLLM 0.22.1 serving Qwen2.5-1.5B-Instruct, against the real Anthropic SDK 1.11.

  • Non-streaming: text returned, usage 23 in / 8 out, stop_reason end_turn, system prompt carried as a top-level field.
  • Streaming envelope in order: message_start, content_block_start, content_block_stop, message_delta, message_stop.
  • Streaming usage: surface reported output_tokens 82; inference_logs recorded completion_tokens 82 for the same request. The first attempt compared the delta count against the surface's own output_tokens, which was circular and passed trivially. Against upstream it did not: 83 deltas for 86 generated tokens. Fixed by requesting stream_options.include_usage.
  • Cancelled mid-stream: returns promptly, no GeneratorExit or unhandled exception in the app log.
  • Image content block: 400, "Unsupported content block type(s): image. Only text is supported."

@Priyanshu-u07
Priyanshu-u07 marked this pull request as ready for review October 3, 2026 16:06
@Priyanshu-u07
Priyanshu-u07 merged commit e831d19 into main Oct 3, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

No Anthropic-compatible API surface

1 participant