Skip to content

Add Dr. GRPO loss normalization - #853

Open
zjn20030811 wants to merge 1 commit into
OpenPipe:mainfrom
zjn20030811:feat/dr-grpo-loss-normalization
Open

Add Dr. GRPO loss normalization#853
zjn20030811 wants to merge 1 commit into
OpenPipe:mainfrom
zjn20030811:feat/dr-grpo-loss-normalization

Conversation

@zjn20030811

Copy link
Copy Markdown

Addresses #364.

Threads TRL-compatible loss_type values (grpo, bnpo, dr_grpo) from public trainer configuration through local, serverless, Unsloth, and Megatron paths. Dr. GRPO uses a fixed B * max_completion_length denominator, including packed completion tails and context-parallel ownership. Incomplete accumulation padding is marked explicitly so dummy microbatches contribute zero while real empty completions retain their batch semantics.

The implicit completion length follows TRL 0.20s 256-token default. Regression tests cover unequal lengths, packed tails, masks, zero-token batches, dummy padding, configuration plumbing, and context-parallel denominator ownership.

Checks: py -3.12 -m pytest -q tests/unit/test_loss_reduction.py src/art/test/test_kl_advantage.py (20 passed); ruff check; ruff format --check; py -3.12 -m compileall. Broader backend suites require Linux/GPU-only dependencies unavailable in the local Windows environment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant