2026-d1-1300-paper-session-b-hot-topic-presentations

Archived conversation · Aug 13, 2026 8:40 PM – Aug 19, 2026 5:37 PM · 57 messages
Thursday, August 13, 2026

Ryan Scherbarth (nvidia) joined the channel
philipp.berdesinski joined the channel
Friday, August 14, 2026

Redfire joined the channel
Saturday, August 15, 2026

darius joined the channel
Sunday, August 16, 2026

Gunethra joined the channel
Monday, August 17, 2026

alnonecat joined the channel
Arulselvan Madhavan joined the channel
Tuesday, August 18, 2026

Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
David Ozog joined the channel
Dan Pitt joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Matthew Fricke joined the channel
Yiltan Temucin joined the channel
Rohit Zambre joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-1300-hot-topic-presentations" to "2026-d1-1300-paper-session-b-hot-topic-presentations"
Wednesday, August 19, 2026

M
Mark Nossokoff 10:53 AM
Is there anything being shared on the Zoom Workspace at the moment? I've been staring at the "will begin soon" slide from "Matthew Fricke's video file" for almost a half hour (since 9:25ish AM Mountain".) It's now 9:50AM and I thought it was supposed to begin at 9:30.
A
Ahmed Khalil (unaffiliated) 10:54 AM
Wondering the same, probably will begin soon.
M
Matthew Fricke 2:47 PM
We open sessions 30 minutes before the start of the day's program. The intro video plays for those 30 minutes to give people time to join Zoom. The first live content was scheduled for 10am MT in the program. Thanks for the question.
B
Bob Lantz 3:08 PM
audio is dropping out
D
darius 3:10 PM
If you have any questions to the current presenter, please write them here.
L
Leila Rashidi 3:13 PM
Is there any interest from hyperscalers or XPU vendors to use XPO for optical scale up? I mean using XPO for XPU!
L
Leila Rashidi 3:25 PM
Is analytical flow level simulation still applicable given latency sensitive AI workload?
B
Bob Lantz 3:26 PM
(as I understood the slide, flow level was the intermediate tier between packet-level and analytic simulation)
1 reply
L
Leila Rashidi 3:32 PM
Agree. I did not get answer. Do you see any benefit from flow level simulation for AI clusters?
L
Leila Rashidi 3:28 PM
Are these results for pkt level simulation? Does simulator benefit from parallelism offered by GPU?
B
Bob Lantz 3:29 PM
is there more information available about your simulator? (I have many questions, including comparing accuracy vs. simulation speed, hearing more about your flow-based and analytic models, etc.)
C
Chris Browning (Black Semi) 3:29 PM
Does this simulation assume or use in-network compute in either configuration? Or is it assuming no BW reduction or collective improvements from in-network compute.
1 reply
L
Laurent Montigny 3:37 PM
Thanks, we are looking at SHARP support in our simulator
B
Bob Lantz 3:31 PM
interested in any public info or papers covering the simulator
4 replies
L
Leila Rashidi 3:32 PM
Already searched and no material is available.
:open_mouth: 1
S
Sayan Ghosh 3:34 PM
Feel free to contact the author directly.
B
Bob Lantz 3:35 PM
maybe we can hear more about it next HotI :D
L
Laurent Montigny 3:36 PM
We will release more info in the next quarter :)
👍 1
L
Leila Rashidi 3:37 PM
What is purpose of sending sent_time in pkt header? This needs per pkt timer at host. So, it does not save state space at host.
1 reply
R
rip.sohan 3:58 PM
It's an architectural trade-off, recall we wanted ot minimize HW changes; some HW does not have this capbility so optional reflection was a good compromise.
L
Leila Rashidi 3:38 PM
Is there any NIC that supports MRC in hardware? Not firmware!
1 reply
R
rip.sohan 3:57 PM
If by HW you mean HW pipeline, AMD does.
P
pgilbert 3:49 PM
what NIC vendors support MRC?
1 reply
R
rip.sohan 3:57 PM
AMD, BRCM, NVIDIA have hardware AFAIK
B
Bob Lantz 3:50 PM
So Rip and Eric, I've worked with you and also with others here at hotI on ultra ethernet, and it seems like there is a lot of overlap between MRC and UE both at the lower levels and even reusing UE's NSCC. Could you talk a bit about the differences, why they matter, and what are the appropriate use cases of MRC vs. UE?
1 reply
R
rip.sohan 4:00 PM
As defined MRC is a strict subset of UE, it does prove some of UE's core principles in the field. There are extensions for reliability and resilience and we're open to feeding them back into UE. I believe the use-cases for both protocols are largely the same, it was a compromise to get something out there quickly and get some field results back. Moving forward I expect UE will dominate but that's a personal opinion.
E
eric.spada 3:53 PM
MRC utilizes the NSCC congestion control algorithm. It also reference the trimming require for NIC and switches
1 reply
B
Bob Lantz 4:01 PM
those are definite overlaps/similarities for sure; interested to hear more about differences, their justification and use cases
T
Taylor Groves 4:02 PM
Thanks for the talk? What are the key differences of YAT compared to other telemetry collection like LDMS or PAPI? Is it the method of collection? Reduced overhead? Or something else?
2 replies
T
Taylor Groves 4:04 PM
Thanks for response!
A
Aaron Welch 5:33 PM
Ah, now I see what was going on with the start of the question - I got tripped up a bit by "YAT" but I'm guessing it was in reference to the "yet another telemetry system" joke from the start, but that was just my silly sense of humour so I wasn't immediately sure what it referred to and just went off context clues. Ha!
O
Omri Mor 4:12 PM
In some of my own research I've found that communication balance over time is an under-discussed topic, so I'm glad to see that others are looking at the problem—especially in the context of message size variability. I've found that often the issue with looking at this is that there's always a tradeoff between gathering fine-grained metrics and overheads—both in memory and storage while the metrics are gathered and later in the complexity of analyzing the data. How was this handled?
1 reply
A
Aaron Welch 5:37 PM
Was your question in reference to my work or that of another?
L
Leila Rashidi 4:17 PM
Does NTT uses scale across for inference? Or remote storage?
L
Leila Rashidi 4:19 PM
What is main motivation of NTT for scale across?
S
Sayan Ghosh 4:21 PM
@Leila Rashidi @Bob Lantz can you please move the questions to the right channel: #C0BQ26A869K
âś… 1
1 reply
B
Bob Lantz 4:23 PM
Oh I didn't realize we weren't still in the afternoon session (or more specifically that the panel had its own channel) - thanks for the heads up
👍 1