Thursday, August 13, 2026
Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026
Ben Michalowicz joined the channel
Sunday, August 16, 2026
Gunethra joined the channel
Monday, August 17, 2026
alnonecat joined the channel
Tuesday, August 18, 2026
Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
David Ozog joined the channel
Sayan Ghosh joined the channel
Dan Pitt joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Matthew Fricke joined the channel
Rohit Zambre joined the channel
Yiltan Temucin joined the channel
aysebilgehan_baspinar joined the channel
Matthew Fricke renamed the channel from "2026-d3-0900-high-performance-and-smart-networking-technologies" to "2026-d3-0900-tutorial-high-performance-and-smart-networking-technologies"
Kapil Shrikhande joined the channel
Subhadeep Bhattacharya joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d3-0900-tutorial-high-performance-and-smart-networking-technologies" to "2026-d3-0900-tutorial-track-3-high-performance-and-smart-networking-technologies"
Friday, August 21, 2026
B
Hi
@channel!! We'll be starting in a few minutes. While I uploaded the latest tutorial slides to dropbox, if any intermediate updates or access issues occur, find the latest slides here:
go.osu.edu/hoti2026-hpn
B
Also please don't hesitate to ask any questions! Dr. Panda and I will be happy to answer them π Seriously, I love doing these tutorials and answering folks' questions for these things
B
Friendly reminder to ask questions if something doesn't make sense (I realize HotI is the place where most of the audience is probably familiar with what's going on hahaha)! Dr. Panda and I are more than happy to answer them regardless π
B
Give a π if you're able to log on and a π if you cannot!
π 3π 1
B
@channel We are taking a 10-minute break! We will be back at 10:40 PT (13:40 ET)
I'll be around to answer questions as they come up π
B
Let me know if you're not able to access the exercise or if something is buggy
B
Where to get the credentials for login and remote IP?
L
Does Sharp supports moe routing (dispatch and combine)?
L
Does it support multi-cast? Sending token to only a subset of connected GPUs?
L
Using software adds delay!
J
Is SHARP still in vogue or becoming a vestigial feature ? like where in the world are ppl turning on this network feature in production cluster?
J
sure its there, but who actually turns it on?
S
@Ben Michalowicz two quesions is any of these switches running Ultra Ethernet on the transport layer ?
S
2) is the multicast IGMP v2 or something else
L
No Ultra Ethernet support
β
1:smiling_face_with_tear: 1
S
Any SmartNic with direct Optical Link Capabilities, aka CPO level of integration as of yet ?
C
What network do you use for MoE? Do you need all to all to pick up experts and broadcast to the gateway/computing systems? Do you allow GPU to dynamically pick up additional expert upon events?
C
Broadcom has smart NIC as well I heard
1 reply
B
They do, yes! I've yet to see whether they have been deployed though. I'm very interested in seeing their performance capabilities as well. Will need to do some searching
B
@Ben Michalowicz tried several logins - cannot access anyway.
6 replies
J
ssh <mailto:ri2tut09@ri2.cse.ohio-state.edu|ri2tut09@ri2.cse.ohio-state.edu> use 09 and use the password listed in table
D
Thanks Jonathan! Good idea.
I've reserved that username for you in the table @BorisKh.
Are you logged in now?
If you type in ssh <mailto:ri2tut09@ri2.cse.ohio-state.edu|ri2tut09@ri2.cse.ohio-state.edu> , then the password, which is in the table, you should be able to log in. I just tested it on my end. Let me know exactly what happens if not...
B
Yes, i logged in,used 09 as Jonathan suggested
π 1
B
Jonathan and David - thanks so much for helping!! π
B
Boris - glad you could log in successfully
C
When is the talk specially address MoE and inference performance?
J
so outside of neoclouds and hyperscalers, and academic experimentation, do you know of anyone who uses smartnics in production cluster academia or other research computing centers
L
What is disadvantages of Cerebras chips vs. Nvidia GPUs?
L
How it compares with groq LPU?
C
@Ben Michalowicz If software stack is a barrier to use CS-* chips, AI can certainly help out to reduce the challenge.
β€οΈ 1
B
I recall that the AISys group at the University of Edinburgh has done extensive work on Cerebras optimizations using CSL, their low-level programming language stack:
arxiv.org/abs/2502.04563
β€οΈ 2
1 reply
B
Thank you for sharing! This is cool to know π
L
Does MTIA 400 also uses NIC chiplet for scale out?
2 replies
B
Hi Leila! Sorry it took this long to get to you:
dl.acm.org/doi/full/10.1145/3695053.3731409 -- Meta's ISCA 2025 paper explains that there is a non-blocking crossbar for routing. Broadly, it uses a switched backplane for scale-out on 72 devices (seen
here)
ai.meta.com/blog/next-generation-meta-training-inference-accelerator⦠-- mentions that they have a NoC architecture for Pe-to-PE (processing element) communication (8x8 on one chip)
This is broadly what I've been able to find in terms of hardware specs for their later generations. Hope this helps!! Your questions were really interesting
L
Where can we find documentation?
S
What sort of confidential computing features are supported in these new NIC developments?
B
They just presented some MTIA info yesterday
β€οΈ 1
L
Yesterdayβs presentation was about MTIA 300
π 1
C
@Dhabaleswar K. Panda Is it possible to add some additional computing accelerating of MAC/GEMM to smart NIC? BTW, do you have a table comparing all the latest AI accelerator chips on the computing & network architecture and performance ?
L
Anybody knows what is topology for MTIA 300 scale up? Does they use switch or full mesh?
2 replies
B
Last I checked, it was a full-mesh topology. Will follow-up later
L
Can you please write name of that NIC vendor?
L
Can you please clarify what does share in performance mean?
L
Is Broadcom Netxtreme proprietary? It is not RoCE?
B
My apologies! I'll be back soon. My computer decided to freeze on me :sweat_smile:
Update: Am back!
B
How would we estimate xCCL performance vs MPI numbers which were shown? I mean if MPI for IB abd RoCE are almost the same, can we expect similar for x CCL?
1 reply
B
The OSU Microbenchmark suite also has benchmarks for xCCL performance! The same performance measurements can be obtained through these benchmarks as the "regular" MPI OMB works for MVAPICH/Open MPI/MPICH/<insert MPI library here>
S
what exactly is size 0 in the previous table?
3 replies
B
0-byte latency! Good for control and synchronization message observation π
S
ok so no payload for size 0 β got it - control only - so in other words, in-band control, akin to higher level control in other protocols, no dedicated control out of band, in band only β could be done with a concurrent paralell network, say if optical at a dedicated Lambda β¦
S
thinking ahead on possible future solutions β¦
B
Very helpful and cool tutorial session! Thx a lot!
S
Very helpful session. Thx
S
extremely good and insightful
B
Thank you all for the great questions! Please feel free to ask any you didn't get a chance to here or through emailing me and Dr. Panda
π 3π 1
B
On a side note, it's really interesting that rumors say TPU v9t is planning to use a 6D torus topology
2 replies
B
I see! Interesting to note. The last time 6D torus was used (to my knowledge) was Supercomputer Fugaku. I'm curious how Google's upcoming TorchTPU integration into PyTorch will leverage them (i.e. from the communication model standpoint), since it's a powerful topology in my opinion, though cable length for wrap-around connections may hinder full connectivity
π 1
B
Yeah Tofu is also something I am thinking about! For TPUs, it's really impressive how effectively static XLA scheduling uses on-chip resources and turn these into tok/s; itβd be also very interesting to look into how they leverage their interconnect.