gspo
Here are 8 public repositories matching this topic...
Reinforcement learning for text generation on MLX (Apple Silicon)
-
Updated
Feb 15, 2026 - Python
Sync vs fully-async agentic RL on verl: multi-turn GRPO, long-tail rollout profiling, staleness ablations — quantifying when async pays off.
-
Updated
Sep 8, 2026 - Python
Medical multi-agent assistant, Qwen3.5-2B LoRA SFT and GSPO experiments, with public datasets and adapters on Hugging Face.
-
Updated
Sep 14, 2026 - Python
OPSG-based test refinement for Java: Stable RL approach to generate maintainable, high-quality unit tests with 98.6% compilation success.
-
Updated
Jan 11, 2026 - Jupyter Notebook
CS336A5 小显存适配:用小显卡直接进行 OLMo-2 全参数强化学习训练,已在单张 RTX 5070 Ti 16GB 上验证。提供低显存训练配置、vLLM 批量生成与复现说明。
-
Updated
Sep 8, 2026 - Python
RL agents across LangChain & LangGraph, 50+ MCP registry tools, PGVector stores, deployment w Dynamo and inference with vLLM & SGLang
-
Updated
Jul 29, 2026 - Python
Add this topic to your repo
To associate your repository with the gspo topic, visit your repo's landing page and select "manage topics."