Thursday, August 13, 2026
Ryan Scherbarth (nvidia) joined the channel
Saturday, August 15, 2026
darius joined the channel
Sunday, August 16, 2026
Gunethra joined the channel
Monday, August 17, 2026
alnonecat joined the channel
Tuesday, August 18, 2026
Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
David Ozog joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Dan Pitt joined the channel
Marc Cohn joined the channel
Matthew Fricke joined the channel
Rohit Zambre joined the channel
Yiltan Temucin joined the channel
aysebilgehan_baspinar joined the channel
Matthew Fricke renamed the channel from "2026-d2-1300-distributed-ai-communication" to "2026-d2-1400-distributed-ai-communication"
Kapil Shrikhande joined the channel
Subhadeep Bhattacharya joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d2-1400-distributed-ai-communication" to "2026-d2-1400-paper-session-c-distributed-ai-communication"
Thursday, August 20, 2026
S
please ask your questions to Session C presenters here, thanks!
M
I am limited on model knowledge so struggling to follow everything here, but would a lager 2k+ scale up domain with 28.8Tbps or 56.7Tbps per GPU package improve the baseline further, especially as we move to 10T or 20T MoE models? What about even larger scale up domains, e.g. up to 5k?
1 reply
B
Sorry for rushing the slides through! Regarding attention mechanisms, I'd like to recommend
htor.inf.ethz.ch/blog/…/lets-stop-calling-everything-linear-attention for a detailed explanation. Per-package bandwidth, yes. The quantity on the critical path is the hot expert's inbound byte count, which is fixed by the model and the batch. Higher per-endpoint bandwidth drains it proportionally faster. Our IB counters say we are nowhere near fabric-limited: peak 82-84 Gbps per node against 4x NDR200 = 800 Gbps installed, with
port_xmit_wait delta 0 in every sampled iteration. Therefore it's endpoint-bound, not fabric-bound, so per-package bandwidth is the right thing to buy.
C
@Bole Ma Excellent talk! We have to predict and pull 1/56 experts ahead of time by the expert gateway, How often do we have to switch to pull additional expert from memory, in case there's a miss? Or is it doable at all dynamically?
1 reply
B
Thank you so much for this fantastic question! I just pulled our measured data and they showed some really interesting patterns: predicting experts works better on some models than others. In models like Qwen3-30B-A3B (GQA), the most-used experts stay fairly predictable from one step to the next, with only about 7.6% of the active experts changing per turn. In models like Qwen3.5-35B-A3B (Qwen switched to Gated DeltaNet Hybrid from Qwen3.5), expert usage shifts rapidly, causing a 72.6% change per turn.
Even when predictions miss, it isn't a total failure, because expert usage in fast-changing models is spread out evenly, picking any set of top experts still covers around 40–50% of your total workload. Missing the exact right expert doesn't hurt performance as much as we might think.
👍 1
L
How network requirements are changing given this imbalance routing and evolution of models?
L
What does buffer occupancy refer to? Buffering at switch may increase delay
1 reply
B
We mean receiver-side occupancy. A hot destination rank has in-flight payloads from all P-1 senders resident simultaneously, so per-endpoint buffer demand scales with fan-in x message size. Sparser routing makes that worse in count, not bytes: more experts per token means more, smaller messages in flight at once.
C
which open source model has the best performance from your experience? or has the most imbalanced network?
1 reply
B
Caveat: Low routing imbalance does not necessarily imply a superior model, and none of the observations below should be taken as a evaluation of overall capability (dense baselines like Qwen3.8-27B remain highly capable).
Interestingly, load imbalance can dynamically shift during training. For example, NVIDIA-Nemotron-3-Nano-30B-A3B initially exhibits high imbalance with concentrated token routing, but naturally rebalances after several hundred steps. Furthermore, for continued pretraining on out-of-domain corpora, hybrid MoE architectures with Gated DeltaNet layers (e.g., Qwen3.5-35B-A3B) appear to be quite robust. However, because they may not yield the optimal aggregate per-rank Gini coefficient, the best architecture ultimately depends on your specific workload requirements.
👍 1
L
How we can test network protocol for all to all? It is not a real all to all! Every token is sent to a subset of GPUs
2 replies
B
I think what matters here is the level of measurement: at the collective level and our operating point, it is dense; but it's not the case for individual token and it can be sparse. In one forward pass, each token goes to k experts, so at most k ranks, but a collective carries a whole batch, and the union covers every destination. Therefore, the primitive is a dense AllToAllv with a skewed count matrix, which is why a uniform AllToAll microbenchmark mis-sizes it. There's one caveat: in very high EP Degree cases (e.g. EP=128) where Local Experts per GPU = 1, we observed that 1.5-4.1% of the ranks might receive nothing during a forward pass, but it constitutes a small fraction.
B
Regarding how to test it, I've recently noticed more efforts targeting the miscalibration issue I mentioned during the presentation. You might be interested in CCL-Bench:
cclbench.ai/index.html
L
Can you please mention a key use-case of federated learning in industry?
L
What is distance for results of slide 10?
L
How LHR compares with ultra Ethernet?
B
Thank you so much for asking these questions! Quite a few of them require me to look into the profile traces, and I will address all of them as soon as possible once I've confirmed the details.
C
Have you consider using OFI instead of UCX. It may be more flexible framework for transports that vary signficantly from a typical cluster.
L
How ACK is taken from critical path?
L
How was congestion control for your scheme?
L
What are applications of NVShmem? any use-case for AI or inference?
1 reply
B
There are many applications using NVSHMEM for both AI inference and HPC. AI examples include deepEP v1 and HPC applications include GROMACS and QUDA
O
Can you make the slides (in Zoom sharing) fullscreen/presentation mode?
L
What does g and p refer to in slide 16?
3 replies
B
they are scalar put and get APIs
B
e.g. nvshmem_int_p puts exactly one integer