Network-faithful simulation of LLM serving and training deployments
-
Updated
Sep 16, 2026 - Python
Network-faithful simulation of LLM serving and training deployments
Early-warning bottleneck profiler for GPU training nodes: GPU + RDMA fabric telemetry with job-level classification
KAI Data Center Builder
Performance metrics for AI / ML cluster
A distributed hardware-software co-design fabric for MoE models (DeepSeek-V3, Mixtral) that eradicates NCCL All-to-All communication stalls via distributed RoCEv2 RDMA virtual address MUX and JAX/XLA SPMD sharding.
Packet-level simulator (ns-3 + RoCEv2/DCQCN/PFC/ECN) comparing fat-tree vs rail-optimized topologies for AI training fabrics. Validated against NCCL on real GPUs.
Ultra-Ethernet compliance test architecture eliminating expensive HBM memory buffers. Design utilizes FPGAs, PRBS payloads, state hash tables, and virtual RDMA.
Visual, exam-focused study guide for Cisco 300-640 DCAI v1.0
A simple experiment applying compressible flow principles to soften distributed gradient communication stalls.
Software examples for RDMA senders and receivers using RDMA WRITE and RDMA READ
PowerShell-based toolkit for Storage Spaces Direct strict bundle deployment, validation, and operational checks in Windows Server 2025 environments.
Lossless-Ethernet benchmark harness for RoCEv2 fabrics: sweeps queue thresholds, ECN marking, and oversubscription to locate throughput collapse and PFC pressure.
Hands-on code, configurations, and visual labs for Cisco 300-640 DCAI v1.0
RoCEv2 control plane simulator — DCQCN congestion control on fat-tree GPU clusters with NCCL Ring AllReduce workloads
GLM-5.3-Flash (320B MoE) at TP=4 across four ASUS Ascent GX10 / DGX Spark nodes on a 100G RoCE fabric: 76 tok/s structured, 3.9M-token fp8 KV pool, 1M context, DFlash2 drafting. Field notes on what broke and why.
To associate your repository with the rocev2 topic, visit your repo's landing page and select "manage topics."