-
Notifications
You must be signed in to change notification settings - Fork 101
Pull requests: InfiniTensor/InfiniLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat: support MiniMax-Text-01 (hybrid Lightning/full-attention MoE, MXFP4, TP/PP)
#563
opened Sep 4, 2026 by
Kritace
Loading…
29 of 34 tasks
feat: lazy-load multimodal package.
#562
opened Sep 4, 2026 by
pengcheng888
Collaborator
Loading…
1 of 49 tasks
feat(nvidia): add reusable GGUF Route B support for Qwen3.5
#559
opened Sep 3, 2026 by
xindongliu594
Loading…
32 of 48 tasks
feat(ascend): enable InfiniOps flash attention
#558
opened Sep 3, 2026 by
baominghelly
Contributor
Loading…
21 of 32 tasks
feat(iluvatar): enable modern Infini stack inference
#557
opened Sep 3, 2026 by
gongchensu
Collaborator
Loading…
4 of 49 tasks
feat(server): add agent support with tool-call and reasoning parsing
#554
opened Sep 1, 2026 by
rubik-hua
Contributor
Loading…
fix: #552 int8 kv cache quantization cannot be enabled
#553
opened Aug 29, 2026 by
BoBoDai
Loading…
38 of 47 tasks
feat: support Ktransformers, CPU-GPU MoE offload via FusedMoE layer
#548
opened Aug 21, 2026 by
whjthu
Contributor
Loading…
37 of 49 tasks
fix(cuda-graph): keep replay metadata dynamic across tensor-parallel ranks
#540
opened Aug 15, 2026 by
junjiewang253-ctrl
Loading…
feat: add aclnnMatmulAllReduce fusion in InfiniLM for Ascend RowParallelLinear
#533
opened Aug 11, 2026 by
ShaneWoof
Contributor
Loading…
feat(hygon): add Qwen3-235B-A3B BF16/W8A8 inference support
#532
opened Aug 7, 2026 by
qinyiqun
Contributor
Loading…
49 tasks
feat(engine): overlap decode steps with asynchronous token handoff
#524
opened Aug 3, 2026 by
qinyiqun
Contributor
Loading…
feat: add Qwen3.6 MoE model support
#521
opened Jul 31, 2026 by
qinyiqun
Contributor
Loading…
49 tasks
perf(server): coalesce streaming SSE output
#517
opened Jul 28, 2026 by
wooway777
Collaborator
Loading…
49 tasks
Previous Next
ProTip!
no:milestone will show everything without a milestone.