2026-d2-0900-keynote-2-networking-innovations-for-gigascale-ai-systems

Archived conversation · Aug 13, 2026 8:41 PM – Aug 21, 2026 7:39 AM · 58 messages
Thursday, August 13, 2026

Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026

Redfire joined the channel
Sunday, August 16, 2026

Gunethra joined the channel
Monday, August 17, 2026

alnonecat joined the channel
Tuesday, August 18, 2026

Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
David Ozog joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Dan Pitt joined the channel
Marc Cohn joined the channel
Matthew Fricke joined the channel
Rohit Zambre joined the channel
Yiltan Temucin joined the channel
Matthew Fricke renamed the channel from "2026-d2-0900-networking-innovations-for-gigascale-ai-systems" to "2026-d2-0900-keynote-networking-innovations-for-gigascale-ai-systems"
aysebilgehan_baspinar joined the channel
Kapil Shrikhande joined the channel
Subhadeep Bhattacharya joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d2-0900-keynote-networking-innovations-for-gigascale-ai-systems" to "2026-d2-0900-keynote-2-networking-innovations-for-gigascale-ai-systems"
Thursday, August 20, 2026

L
Leila Rashidi 11:04 AM
Q: where CPU racks are deployed? Front-end or backend?
D
Dan Pitt 11:07 AM
When we were first promoting SDN in ONF we predicted that networking would become an arm of computing. Gilad’s presentation shows that this has come to pass.
C
Chris Browning (Black Semi) 11:07 AM
CPU Racks for what purpose? The ones working with the GPUs are integrated with the GPU on the same board.
L
Leila Rashidi 11:12 AM
What is the maximum distance supported by scale across solution? Has scale across solution been deployed in production? Is scale across solution based on back-to-sender notification?
1 reply
S
scots 11:25 AM
Scale-across begins roughly around 500 meters; but can span hundreds of kilometers. Spectrum-XGS is based on Spectrum-X Ethernet with distance-aware algorithms
πŸ‘ 1
L
Leila Rashidi 11:14 AM
How Enfabrica IP help Nvidia regarding KV cache or storage?
Y
Yusuke Ohara(Self) 11:15 AM
NVLink has now been expanded to two levels, allowing eight NVL72s to be connected within a single NVLink domain. But how are the individual NVL72s connected to one another? Is it correct to assume that half of the links from the NVSwitch are used to connect the switches to each other rather than to the GPUs?
2 replies
S
scots 11:43 AM
What I know what was disclosed about this topic is here; and a good read πŸ™‚ NVIDIA Vera Rubin POD: Seven Chips, Five Rack-Scale Systems, One AI Supercomputer | NVIDIA Technical Blog
Y
Yusuke Ohara(Self) 11:57 AM
Yes, I’ve read this, but I can’t find any explanation of how the nine NVSwitch trays housed in each MGX rack are connected to one another.
Looking at the photos of the NVL576 prototype, it appears that two of the four NVSpines have been omitted, and instead, optical cables have been routed via media converters to connect to the 16 L2 NVSwitches housed in the rack to the right of the eight NVL72 units. Is this interpretation correct?
L
Leila Rashidi 11:26 AM
What is Nvidia position about XPO ?
A
ashkan.sobhani 11:38 AM
Does NVIDIA consider NIC disaggregation in its roadmap for NICs?
L
Leila Rashidi 11:43 AM
Does distance aware load balancing refer to choosing shortest path? Or it is something beyond this?
πŸ‘ 1
1 reply
S
scots 11:47 AM
Does not always mean shortest path. It depends on configuration of the scale-across. Spectrum-XGS is based on algorithms to ensure that you are not paying additional latency penalty; like you do with OTS deep buffer switching.
πŸ‘ 1
A
ashkan.sobhani 11:44 AM
Does distance-aware load balancing spray packets across DCIs?
K
Kazuaki Ueda 11:44 AM
In yesterday's Meta keynote, it was highlighted that the "Scale-across" domain involves highly heterogeneous environments, particularly regarding distance (propagation delay) and bandwidth. Given this context, can Spectrum-XGS consistently guarantee optimal performance under such diverse conditions? I would love to hear your thoughts on the future challenges in this area!
C
Cissy Yuan 11:45 AM
Is DCI or context memory system also aware of type of network access, prefill vs decoding, normal KV $ vs MOE?
D
Dan Pitt 11:45 AM
What is the relation between DOCA and DSX?
A
andrewliang823 11:46 AM
you talked about Oberon Rack, how about the Kyber Rack for 144 per Rack?
N
nicky 11:46 AM
What are the primary challenges you faced with hot liquid cooling especially with the large copper backplane?
L
Leila Rashidi 11:46 AM
Why SSD is not located in GPU server?
1 reply
J
jfkim 11:59 AM
The GPU servers typically do have SSDs for boot, OS, logs, etc. But in most cases they don't have enough capacity or endurance for AI data or KV cache.
G
gwhelan 11:48 AM
What about the Out-of-Band Management network? What's next after 10G?
M
Mike Capuano 11:48 AM
Will you deploy separate dedicated Vera racks for agentic services? If so, what is the bandwidth per GPU?
T
Thomas Deucher 11:49 AM
How will you approach scaling bandwidth with the Spectrum-X switch in future generations? WDM?
D
Danna Wang 11:49 AM
How do we access the DSX Air?
2 replies
S
scots 11:56 AM
Start here: NVIDIA DSX Air Platform for AI Factory Simulation ; scroll down towards the end of the page to take a test drive
πŸ‘ 2
D
Danna Wang 12:09 PM
Amazing, thanks for sharing this @scots !
N
Nathan Brown 11:51 AM
How does MRC (or UEC) factor into the Scale-Out or Scale-Across RDMA network design?
M
Mohammed Mahfuz 11:55 AM
Hi, Shainer.
Thanks for the nice talk
I have a question about Nvidia Spectrum CPO: What are the energy efficiency and bandwidth density?
R
Ramtin Soleymani 11:57 AM
For a third-party XPU connecting through NVLink Fusion, what tends to be the hardest part to get right in hardware?
J
Jan Gray 11:58 AM
How might emerging new memory tiers, such as HBF, high bandwidth flash, impact CMX context storage architecture and its interconnect requirements?
(Fantastic talk reviewing your stunning chiplet-to-datacenters co-design, thank you.)
R
Rabindra Guha (Cerio) 12:02 PM
is Ravi sharing?
πŸ‘ 1βœ… 1
M
Matthew Fricke 12:07 PM
Thanks for letting us know
R
Rabindra Guha (Cerio) 12:15 PM
@Ravi Mahatme What is the bandwidth overhead to support cache coherency vs using just CXL.io, for shared memory?
2 replies
R
rmahatme363 12:27 PM
We support cxl.io & cxl.mem . There is no native cache coherency
R
Rabindra Guha (Cerio) 12:32 PM
Then why CXL. Why not just have a PCIe Gen6 <->Memory (HBM/DDR) bridge? Are you just using the Port Routing feature of the CXL protocol?
Friday, August 21, 2026

G
gwhelan 7:39 AM
@shainer Great keynote. Super insightful. Would you, or a colleague, be interested in a short chat about our "Adaptive Wireless Infrastructure for AI Data Centers"? It's a great next generation Out-of-Band Management network. 100G, Ethernet Compatible and no more wires. Truly out-of-band. Teradio is a spinout company from Northeastern Univ. (NU) and we've just raised $5M from the US NSF for advanced research on AI DCs wireless networking. Nvidia is a sponsor of the research, though we'd like to have a strategic discussion with you as well. Thank you, Greg Whelan, CEO Teradio, Inc. <mailto:gwhelan@teradop.co|gwhelan@terad>io.co