2026-d3-0900-tutorial-track-3-high-performance-and-smart-networking-technologies

Archived conversation · Aug 13, 2026 8:44 PM – Aug 21, 2026 6:39 PM · 91 messages
Thursday, August 13, 2026

Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026

Ben Michalowicz joined the channel
Sunday, August 16, 2026

Gunethra joined the channel
Monday, August 17, 2026

alnonecat joined the channel
Tuesday, August 18, 2026

Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
David Ozog joined the channel
Sayan Ghosh joined the channel
Dan Pitt joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Matthew Fricke joined the channel
Rohit Zambre joined the channel
Yiltan Temucin joined the channel
aysebilgehan_baspinar joined the channel
Matthew Fricke renamed the channel from "2026-d3-0900-high-performance-and-smart-networking-technologies" to "2026-d3-0900-tutorial-high-performance-and-smart-networking-technologies"
Kapil Shrikhande joined the channel
Subhadeep Bhattacharya joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d3-0900-tutorial-high-performance-and-smart-networking-technologies" to "2026-d3-0900-tutorial-track-3-high-performance-and-smart-networking-technologies"
Friday, August 21, 2026

B
Ben Michalowicz 10:58 AM
Hi @channel!! We'll be starting in a few minutes. While I uploaded the latest tutorial slides to dropbox, if any intermediate updates or access issues occur, find the latest slides here: go.osu.edu/hoti2026-hpn
B
Ben Michalowicz 10:59 AM
Also please don't hesitate to ask any questions! Dr. Panda and I will be happy to answer them πŸ™‚ Seriously, I love doing these tutorials and answering folks' questions for these things
B
Ben Michalowicz 11:30 AM
Friendly reminder to ask questions if something doesn't make sense (I realize HotI is the place where most of the audience is probably familiar with what's going on hahaha)! Dr. Panda and I are more than happy to answer them regardless πŸ™‚
J
Jonathan Skone 11:57 AM
Is RoCE sunsetting?
B
Ben Michalowicz 12:00 PM
Give a πŸ‘ if you're able to log on and a πŸ‘Ž if you cannot!
πŸ‘ 3πŸ‘Ž 1
B
Ben Michalowicz 12:01 PM
πŸ‘ 2
B
Ben Michalowicz 12:30 PM
@channel We are taking a 10-minute break! We will be back at 10:40 PT (13:40 ET)

I'll be around to answer questions as they come up πŸ™‚
B
Ben Michalowicz 12:31 PM
Once again, the updated slides are here: go.osu.edu/hoti2026-hpn
The github link for the hands on is here: github.com/OSU-Nowlab/HPN-Tutorial
B
Ben Michalowicz 12:54 PM
Let me know if you're not able to access the exercise or if something is buggy
B
BorisKh 12:58 PM
Where to get the credentials for login and remote IP?
J
L
Leila Rashidi 12:59 PM
Does Sharp supports moe routing (dispatch and combine)?
L
Leila Rashidi 1:01 PM
Does it support multi-cast? Sending token to only a subset of connected GPUs?
L
Leila Rashidi 1:02 PM
Using software adds delay!
J
Jonathan Skone 1:03 PM
Is SHARP still in vogue or becoming a vestigial feature ? like where in the world are ppl turning on this network feature in production cluster?
J
Jonathan Skone 1:05 PM
sure its there, but who actually turns it on?
S
skarevik 1:05 PM
@Ben Michalowicz two quesions is any of these switches running Ultra Ethernet on the transport layer ?
S
skarevik 1:06 PM
2) is the multicast IGMP v2 or something else
S
skarevik 1:06 PM
correction IGMP V3
L
Leila Rashidi 1:07 PM
No Ultra Ethernet support
βœ… 1:smiling_face_with_tear: 1
S
skarevik 1:16 PM
Any SmartNic with direct Optical Link Capabilities, aka CPO level of integration as of yet ?
C
Cissy Yuan 1:16 PM
What network do you use for MoE? Do you need all to all to pick up experts and broadcast to the gateway/computing systems? Do you allow GPU to dynamically pick up additional expert upon events?
C
Cissy Yuan 1:16 PM
Broadcom has smart NIC as well I heard
1 reply
B
Ben Michalowicz 1:18 PM
They do, yes! I've yet to see whether they have been deployed though. I'm very interested in seeing their performance capabilities as well. Will need to do some searching
B
BorisKh 1:17 PM
@Ben Michalowicz tried several logins - cannot access anyway.
6 replies
J
Jonathan Skone 1:19 PM
ssh <mailto:ri2tut09@ri2.cse.ohio-state.edu|ri2tut09@ri2.cse.ohio-state.edu> use 09 and use the password listed in table
B
BorisKh 1:21 PM
Thx!
D
David Ozog 1:22 PM
Thanks Jonathan! Good idea.

I've reserved that username for you in the table @BorisKh.

Are you logged in now?

If you type in ssh <mailto:ri2tut09@ri2.cse.ohio-state.edu|ri2tut09@ri2.cse.ohio-state.edu> , then the password, which is in the table, you should be able to log in. I just tested it on my end. Let me know exactly what happens if not...
B
BorisKh 1:23 PM
Yes, i logged in,used 09 as Jonathan suggested
πŸ‘ 1
B
Ben Michalowicz 1:25 PM
Jonathan and David - thanks so much for helping!! πŸ˜„
B
Ben Michalowicz 1:25 PM
Boris - glad you could log in successfully
C
Cissy Yuan 1:18 PM
When is the talk specially address MoE and inference performance?
J
Jonathan Skone 1:20 PM
so outside of neoclouds and hyperscalers, and academic experimentation, do you know of anyone who uses smartnics in production cluster academia or other research computing centers
L
Leila Rashidi 1:21 PM
What is disadvantages of Cerebras chips vs. Nvidia GPUs?
L
Leila Rashidi 1:26 PM
How it compares with groq LPU?
C
Cissy Yuan 1:26 PM
@Ben Michalowicz If software stack is a barrier to use CS-* chips, AI can certainly help out to reduce the challenge.
❀️ 1
B
Bole Ma 1:30 PM
I recall that the AISys group at the University of Edinburgh has done extensive work on Cerebras optimizations using CSL, their low-level programming language stack: arxiv.org/abs/2502.04563
❀️ 2
1 reply
B
Ben Michalowicz 1:51 PM
Thank you for sharing! This is cool to know πŸ™‚
L
Leila Rashidi 1:32 PM
Does MTIA 400 also uses NIC chiplet for scale out?
2 replies
B
Ben Michalowicz 2:09 PM
Hi Leila! Sorry it took this long to get to you:
dl.acm.org/doi/full/10.1145/3695053.3731409 -- Meta's ISCA 2025 paper explains that there is a non-blocking crossbar for routing. Broadly, it uses a switched backplane for scale-out on 72 devices (seen here)

ai.meta.com/blog/next-generation-meta-training-inference-accelerator… -- mentions that they have a NoC architecture for Pe-to-PE (processing element) communication (8x8 on one chip)

This is broadly what I've been able to find in terms of hardware specs for their later generations. Hope this helps!! Your questions were really interesting
L
Leila Rashidi 6:39 PM
Thanks
L
Leila Rashidi 1:32 PM
Where can we find documentation?
S
sankar ramamoorthi 1:33 PM
What sort of confidential computing features are supported in these new NIC developments?
B
BorisKh 1:33 PM
They just presented some MTIA info yesterday
❀️ 1
L
Leila Rashidi 1:34 PM
Yesterday’s presentation was about MTIA 300
πŸ‘ 1
C
Cissy Yuan 1:35 PM
@Dhabaleswar K. Panda Is it possible to add some additional computing accelerating of MAC/GEMM to smart NIC? BTW, do you have a table comparing all the latest AI accelerator chips on the computing &amp; network architecture and performance ?
L
Leila Rashidi 1:35 PM
Anybody knows what is topology for MTIA 300 scale up? Does they use switch or full mesh?
2 replies
B
Ben Michalowicz 1:36 PM
Last I checked, it was a full-mesh topology. Will follow-up later
L
Leila Rashidi 1:37 PM
Thanks
L
Leila Rashidi 1:36 PM
Can you please write name of that NIC vendor?
B
Ben Michalowicz 1:36 PM
Chelsio
B
Ben Michalowicz 1:37 PM
L
Leila Rashidi 1:39 PM
Can you please clarify what does share in performance mean?
L
Leila Rashidi 1:42 PM
Is Broadcom Netxtreme proprietary? It is not RoCE?
C
Cissy Yuan 1:45 PM
yes we can
B
Ben Michalowicz 1:46 PM
My apologies! I'll be back soon. My computer decided to freeze on me :sweat_smile:
Update: Am back!
B
BorisKh 1:55 PM
How would we estimate xCCL performance vs MPI numbers which were shown? I mean if MPI for IB abd RoCE are almost the same, can we expect similar for x CCL?
1 reply
B
Ben Michalowicz 2:01 PM
The OSU Microbenchmark suite also has benchmarks for xCCL performance! The same performance measurements can be obtained through these benchmarks as the "regular" MPI OMB works for MVAPICH/Open MPI/MPICH/&lt;insert MPI library here&gt;
S
skarevik 1:56 PM
what exactly is size 0 in the previous table?
3 replies
B
Ben Michalowicz 1:59 PM
0-byte latency! Good for control and synchronization message observation πŸ™‚
S
skarevik 2:01 PM
ok so no payload for size 0 β€” got it - control only - so in other words, in-band control, akin to higher level control in other protocols, no dedicated control out of band, in band only β€” could be done with a concurrent paralell network, say if optical at a dedicated Lambda …
S
skarevik 2:01 PM
thinking ahead on possible future solutions …
B
BorisKh 2:03 PM
Very helpful and cool tutorial session! Thx a lot!
S
sankar ramamoorthi 2:03 PM
Very helpful session. Thx
S
skarevik 2:03 PM
extremely good and insightful
B
Ben Michalowicz 2:05 PM
Thank you all for the great questions! Please feel free to ask any you didn't get a chance to here or through emailing me and Dr. Panda
πŸ™Œ 3πŸ‘ 1
B
Ben Michalowicz 2:06 PM
Y'all were amazing πŸ™‚
B
Bole Ma 4:02 PM
On a side note, it's really interesting that rumors say TPU v9t is planning to use a 6D torus topology
2 replies
B
Ben Michalowicz 4:04 PM
I see! Interesting to note. The last time 6D torus was used (to my knowledge) was Supercomputer Fugaku. I'm curious how Google's upcoming TorchTPU integration into PyTorch will leverage them (i.e. from the communication model standpoint), since it's a powerful topology in my opinion, though cable length for wrap-around connections may hinder full connectivity
πŸ‘ 1
B
Bole Ma 4:11 PM
Yeah Tofu is also something I am thinking about! For TPUs, it's really impressive how effectively static XLA scheduling uses on-chip resources and turn these into tok/s; it’d be also very interesting to look into how they leverage their interconnect.