Initialize Miles LoRA A matrices with kaiming-uniform like the megatron backend - #10
Draft
kevintli wants to merge 1 commit into
Conversation
|
I'll fix CI failures and address comments from users with write access that start with 'Devin'.
|
devin-ai-integration
Bot
force-pushed
the
devin/1790783613-miles-lora-kaiming-init
branch
from
September 30, 2026 20:25
7f60417 to
216ab5a
Compare
…on backend Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
devin-ai-integration
Bot
force-pushed
the
devin/1790783613-miles-lora-kaiming-init
branch
from
September 30, 2026 20:48
216ab5a to
0ac5751
Compare
kevintli
added this pull request to stack #13
September 30, 2026 22:15
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Miles backend never sets
args.lora_A_init_method, so Miles'create_multi_lora_instancefalls back togetattr(args, "lora_A_init_method", "xavier"). In the bundled Megatron-BridgeMultiLoRAlayers,"xavier"meansxavier_normal_. Spindle's own megatron LoRA backend (megatron_runtime/lora/model.py) passeslora_A_init_method="kaiming"(kaiming_uniform_(a=sqrt(5)), the PEFT default). Thinking Machines' "LoRA Without Regret" post says Tinker uses that PEFT parametrization: uniform A with scale1/sqrt(d_in), zero B, alpha 32.Why it matters: B starts at zero, and Adam's per-step update to B is about
lrregardless of scale. So the early weight changeΔW = (α/r)·ΔB·Ascales with std(A), and xavier gives a higher effective LR than Tinker at the samelearning_rate.Measured std(A) in the exported step-0 adapters for gpt-oss-20b (layer 3 and lm_head). The PEFT reference is
1/sqrt(3·d_in):This PR fixes the dense q/k/v/lm_head targets exactly. Two gaps remain because the init runs on local, packed tensors, which gives the wrong fan-in:
o_projsees the 512-wide TP shard.[local_experts, rank, in]seerank*in.Fixing those needs per-expert, global-fan-in init in Megatron-Bridge. That is out of scope here; this PR only makes the Miles path consistent with Spindle's megatron path.
Found while reproducing jasper-lu/sec-search-rl, where Spindle policies collapsed to short trajectories faster than the hosted-Tinker reference.
Link to Devin session: https://modal.devinenterprise.com/sessions/f53cfabb210146de8f0338fe388d7973
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/f53cfabb210146de8f0338fe388d7973?variant=devin
Requested by: @kevintli