feat: serve the Anthropic Messages API on /v1/messages - #372
Merged
Merged
Conversation
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
Collaborator
Author
|
Verified on a live deployment, g5.xlarge (A10G), kind cluster via deploy/gpu-kind.sh, vLLM 0.22.1 serving Qwen2.5-1.5B-Instruct, against the real Anthropic SDK 1.11.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Anything written against the Anthropic Messages API, the Claude SDK and the agent frameworks built on it, could not point at the gateway at all, because only OpenAI-shaped routes existed.
/v1/messagestranslates the request into the OpenAI shape used internally,runs the ordinary completion path, and translates the answer back. Nothing in that path is aware of the surface, so API key resolution, quotas, rate limiting, routing and logging are the same code rather than a second copy. The diff on existing files is 39 lines: a route, an export and an alias.A surface is the mirror of a provider adapter. One translates between a client's shape and ours, the other between an upstream's shape and ours:
#368 built the right-hand side. This is the left.
Three things worth a look in review
x-api-keyfalls back to the bearer header, because Anthropic clients send the key that way and would otherwise fail to authenticate. Sandbox is deliberately excluded from that fallback: it carries a JWT thatextract_api_keyverifies, and x-api-key must not be a way around it.The stream translator closes its envelope in a
finally. Anthropic wraps text in an explicit envelope that OpenAI carries implicitly, so most of it is synthesised rather than translated — and if upstream dies mid-stream a client is otherwise left waiting onmessage_stopforever.Translator state lives on the instance here, unlike #368's, because a surface translator is built per request where provider adapters are cached and shared. Covered by a concurrent-streams test.
Known gaps
Only text content is supported. Image and tool blocks are rejected with a 400 naming the type, rather than dropped, so a request never quietly loses content and answers anyway.
toolsis not translated. Largely moot today, since vLLM deployments launch without--enable-auto-tool-choiceand so cannot emit structured tool calls either way./v1/messages/count_tokensis not routed. vLLM serves it; the SDK onlycalls it when asked to.Streamed
usage.output_tokenscounts deltas, not tokens. Upstream reports an exact count only whenstream_options.include_usageis requested, which the surface does not control. One delta is one token on vLLM and anapproximation anywhere that batches.
Why this is a draft
Everything upstream of the route is mocked. The translation is verified by the real Anthropic SDK, but no request in these tests reaches vLLM. Draft until it has been run against a live deployment.
One specific thing to check there: cancel a stream mid-response. The tests cover upstream dying, but a client disconnecting is different, Starlette cancels the generator rather than raising into it, and whether the
finallystill closes the envelope under that cancellation is the one claim here no mock can settle.Closes #369