feat(iluvatar): enable modern Infini stack inference - #557
Open
gongchensu wants to merge 4 commits into
Open
Conversation
voltjia
requested changes
Sep 4, 2026
Comment on lines
+17
to
+26
| std::size_t implementation_index_for_device( | ||
| infini::ops::Device::Type device_type) { | ||
| if (device_type == infini::ops::Device::Type::kIluvatar) { | ||
| return 0; | ||
| } | ||
| if (device_type == infini::ops::Device::Type::kMoore) { | ||
| return 8; | ||
| } | ||
| return 16; | ||
| } |
Collaborator
There was a problem hiding this comment.
这块加个 TODO 吧,感觉老师应该不希望 InfiniLM 里面有这种东西 hhh,后面看看怎么办吧。
Collaborator
There was a problem hiding this comment.
这个文件可以本地留存一份,但是先不放到 InfiniLM 仓库里。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
scripts/build_infini_stack.py, including InfiniRT, InfiniOps, and InfiniCCL configuration.Motivation
The
refactor/adopt-modern-infini-stackbranch requires platform integrations to use the standalone InfiniRT, InfiniOps, and InfiniCCL components instead of the legacy implementations previously contained in InfiniCore.Before this change, InfiniLM did not accept
iluvataras a model runner device, the stack builder could not configure the Iluvatar toolchain, and the attention adapters did not register the canonical InfiniOps Iluvatar providers.This change enables the 9G-8B model on Iluvatar BI-V150 through static attention, paged attention with InfiniRT graph execution, and explicit FlashAttention with InfiniRT graph execution.
N/A - no linked issue.
Type of Change
feat— new feature / new modelfix— bug fixperf— performance improvement (no behavioral change)refactor— code restructuring without behavior changetest— adding or fixing tests onlydocs— documentation onlybuild/ci— build system or CI configurationchore— tooling, formatting, or other non-code changesTest Results of Involved Models on Supported Platforms (Please attach screenshots)
Benchmark / Performance Impact
Notes for Reviewers
CI / ChatOps
Checklist
Title, Branch, and Commits
feat(nvidia): …,fix(cuda/gemm): …).<type>/xxx-yyyy-zzzzwhere<type>matches the PR title's Conventional Commits type and words are joined with hyphens (seeCONTRIBUTING.md§Branches).CONTRIBUTING.md§Pull Requests).main— the branch is rebased cleanly on top of the currentmain.fixup!/squash!/wipcommits remain.Scope and Design
CONTRIBUTING.md§Code/General).printf/std::cout/print(...)left behind, orTODOwithout an owner and issue link.General Code Hygiene (applies to all languages)
CONTRIBUTING.md§Code/General).CONTRIBUTING.md§Code/General).the `seqlens_k` tensor) (CONTRIBUTING.md§Code/General).CONTRIBUTING.md§Code/General).CONTRIBUTING.md§Code/General; §Python).C++ Specific (if C++ files changed)
CONTRIBUTING.md§C++).CONTRIBUTING.md§C++).new/delete; RAII / smart pointers / existing allocators are used.scripts/format.py.csrc/models/llama_legacy/.Python Specific (if Python files changed)
CONTRIBUTING.md§Python).CONTRIBUTING.md§Python).scripts/format.py.python/infinilm/auto_config.py.Testing
examples/test_infer.py), or specify the reason for skipping.examples/bench.py), or specify the reason for skipping.test/bench/test_benchmark.py), or specify the reason for skipping.python/infinilm/server/inference_server.py+scripts/test_perf.py), or specify the reason for skipping.Build, CI, and Tooling
/retestwas requested.Documentation
README.md,CONTRIBUTING.md, or inline docs updated when behavior, build flags, or developer workflow changed.!orBREAKING CHANGE:footer.Security and Safety