Skip to content

Serve gpt-oss LoRA from BF16 weights with the Triton MoE runner - #7

Draft
kevintli wants to merge 1 commit into
devin/1790748858-gpt-oss-lm-headfrom
devin/1790751602-gpt-oss-bf16-sampler
Draft

kevintli wants to merge 1 commit into
devin/1790748858-gpt-oss-lm-headfrom
devin/1790751602-gpt-oss-bf16-sampler

Conversation

@kevintli

Copy link
Copy Markdown

Summary

With the gpt-oss preset, the LoRA sampler failed at startup:

sglang/srt/lora/layers.py: NotImplementedError: LoRA MoE not supported for backend MoeRunnerBackend.TRITON_KERNELS

openai/gpt-oss-20b ships MXFP4 experts, and SGLang routes those through triton_kernel. SGLang's LoRA MoE layer only supports triton (which needs get_triton_quant_info, and Mxfp4MoEMethod doesn't have it) or marlin (CompressedTensors/NvFp4 only). Miles has the same constraint and says so: examples/lora/run-gpt-oss-20B-megatron-moe-lora.sh has "need to use bf16 ckpt when enable triton moe backend, eg, lmsys/gpt-oss-20b-bf16", and its gpt-oss LoRA CI uses gpt-oss-20b-bf16.

Changes:

  • New recipe field model_weights: the HF repo for the base weights, defaulting to model. The served/client-facing base_model is still model; only the downloaded assets change.
    DeploymentConfig.weights_repo = recipe.model_weights or recipe.model
    asset_path = f"/assets/{weights_repo}"          # trainer hf_checkpoint + SGLang --model-path
    prepare_model_assets: snapshot_download(repo_id=weights_repo, ...)
    model_weights is now part of trainer_identity() and inference_identity(), so switching weights redeploys both.
  • gpt-oss preset: model_weights = "lmsys/gpt-oss-20b-bf16" (the MXFP4 release upcast to BF16, same architecture and config apart from quantization_config), plus sglang_cfg.moe_runner_backend = "triton".

Stacked on #6.

Link to Devin session: https://modal.devinenterprise.com/sessions/f53cfabb210146de8f0338fe388d7973
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/f53cfabb210146de8f0338fe388d7973?variant=devin
Requested by: @kevintli

@devin-ai-integration

Copy link
Copy Markdown

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1790751602-gpt-oss-bf16-sampler branch from 0c3a2ef to 07fca4a Compare September 30, 2026 20:48
@kevintli
kevintli added this pull request to stack #13 September 30, 2026 22:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant