fix: echo back the model name the client sent - #367
Merged
Merged
Conversation
Signed-off-by: Priyanshu-u07 <connect.priyanshu8271@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Request
{"model": "vllm-14b"}came back as"model": "Qwen/Qwen2.5-14B-Instruct-AWQ". Every response path that carries a model field now echoes the name the client used.Non-streaming
One shared helper,
echo_requested_modelinworker_routing.py, alongsideupstream_modelwhich already owns the send-side naming. All four handlers call it: chat completions, embeddings, images and video. It is a no-op for engines whose response has no model field, so images and video are covered without assuming what a custom diffusion engine returns.The client's name is still available at that point — the pipeline only ever substitutes the upstream id into the outgoing body, never onto
ctx.model.Streaming
This needed more than an overwrite. Upstream is read with
aiter_raw, so chunks are arbitrary byte fragments: a singledata:line, or the model name inside it, can arrive split across two chunks, and a chunk cannot be rewritten on its own.process_streamnow re-frames into whole lines when a rewrite is requested, holding any trailing fragment for the next chunk and flushing at the end. It also decodes incrementally, because a per-chunkdecode(errors="ignore")throws away both halves of a multibyte character that straddles a boundary and deletes it from the response. A 3-byte Devanagari character split after its first byte disappeared entirely before that fix.Only events actually being rewritten are re-serialised. Error payloads,
[DONE], blank separators and non-JSON lines are returned byte for byte. With no rewrite requested the originalyield chunkpath runs unchanged, so other callers are unaffected, and the usage-parsing buffer is untouched.Tests
25 tests. The ones that matter are the boundary cases: a line split across chunks, a split inside the model name, one byte at a time, a multibyte character split mid-sequence, and a final line with no trailing newline.
Each fix is mutation-verified — a naive byte replace instead of line framing passes 11 tests and fails exactly the 3 boundary ones; the old per-chunk decode fails the multibyte test; removing the helper call fails that path's echo test.
Closes #353