2026-d1-1040-paper-session-a-network-design-at-scale

Archived conversation · Aug 13, 2026 8:39 PM – Aug 19, 2026 2:09 PM · 66 messages
Thursday, August 13, 2026

Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026

Arulselvan Madhavan joined the channel
Redfire joined the channel
Sunday, August 16, 2026

Gunethra joined the channel
Monday, August 17, 2026

alnonecat joined the channel
Tuesday, August 18, 2026

Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
David Ozog joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Dan Pitt joined the channel
Matthew Fricke joined the channel
Yiltan Temucin joined the channel
Rohit Zambre joined the channel
aysebilgehan_baspinar joined the channel
Kapil Shrikhande joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-1040-network-design-at-scale" to "2026-d1-1040-paper-network-design-at-scale"
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-1040-paper-network-design-at-scale" to "2026-d1-1040-paper-session-a-network-design-at-scale"
Wednesday, August 19, 2026

K
Kapil Shrikhande 12:39 PM
Hi everyone !
👏 1
K
Kapil Shrikhande 12:43 PM
Welcome to the session : Network Design and Scale.
K
Kapil Shrikhande 12:43 PM
Feel free to enter your questions here.
M
Mike Capuano 12:58 PM
Do you see MRLS as applicable to back-end scale out networks? It seems like increasing diameter beyond 2 would not be desirable and that MRLS would only be applicable to the front end network.
R
Rabindra Guha (Cerio) 12:59 PM
This sound similar to Rockport Networks Torus Networking technology.
1 reply
L
Leila Rashidi 1:07 PM
Is this posted in website of Cerio?
P
pgilbert 12:59 PM
Does this use stndard routing protocols or is there a new routing protocol?
4 replies
R
Rabindra Guha (Cerio) 1:02 PM
I think any packet based protocol, that has a SRC and DST address. Thus Ethernet is a supported protocol.
👍 1
P
pgilbert 1:02 PM
he mentioned you need to do 2 lookup's that new?
P
pgilbert 1:05 PM
still not clear 😀
A
Alejandro Cano 2:09 PM
The underlying minimal routing tables can be built using standard mechanisms, but what changes is how packets are forwarded. Usually, standard routings do 1 lookup to the routing table. Polarized routing needs 2 lookups of the minimal routing table: one querying the source and the other for the destination. It requires both to compute the correct ports.
K
Kapil Shrikhande 1:00 PM
Does your cost model take into account differences in copper vs. optical interconnect. And how does the MRLS topology compare to Fat-tree in its breakdown of optical vs. copper links .
S
Sayan Ghosh 1:01 PM
Is it possible to do parallel simulation in CAMINOS, because with more endpoints, serial simulation overhead can be significant? Also, have you considered including the topology/routing on existing discrete event simulators like SST/Macro (which may support parallel simulation)?
L
Leila Rashidi 1:02 PM
What is hardware requirement in switch?
C
Cissy Yuan 1:07 PM
@Alejandro Cano So MRLS expands the network to new nodes with the trade off to reduce the bandwidth of existing nodes? Is it possible to add another layer of fat tree at top when the workload is too much, so it's 2D MRLS?
L
Leila Rashidi 1:07 PM
Does it run over commodity switches?
L
Leila Rashidi 1:09 PM
What about per pkt state?
K
Kapil Shrikhande 1:32 PM
Please enter your questions for the second Paper of the session.
L
Laurent Montigny 1:33 PM
Any insight regarding the impact on power for this study with optical?
3 replies
T
Taylor Groves 1:37 PM
4.3pJ/b used based of hardware from the lab. That's all in SerDes, laser, optics.
T
Taylor Groves 1:38 PM
You'd compare that against your long-reach 224G SerDes from your favorite source for the electrical link. If it's a pluggable optics its going to be much higher ~16pJ/b.
👍 1
L
Laurent Montigny 1:41 PM
Great, thank for the details
C
Chris Browning (Black Semi) 1:33 PM
Do you think some of the assumptions in the analysis will change with the introduction of XPO formfactore, increasing the density of pluggable optics for a more classic system (copper/optics)
1 reply
T
Taylor Groves 1:39 PM
XPO is a good density solution keep pluggables relevant for a while longer, but not going to be as efficient as CPO.
K
Kapil Shrikhande 1:34 PM
Is there a reason you showed the scale-up domain up to 1152 GPUs ? Where is the constraint coming from – is it from the optics, o# racks, distance ?
1 reply
M
Mike Capuano 1:36 PM
It is limited by switch radix
👍 1
N
Nitin Garg 1:35 PM
Can someone share the link to the paper?
B
Bole Ma 1:35 PM
Do you need more fault-tolerant designs for photonic interconnects?
1 reply
T
Taylor Groves 1:43 PM
You do add more components (e.g. laser + PIC) with photonics, but there are lots of options for making these reliable and replaceable. External laser, detatchable fibers, redundancy in design, etc.
C
Chris Browning (Black Semi) 1:36 PM
The papers are up in the conference proceedings
S
Sayan Ghosh 1:36 PM
@Nitin Garg paper can be fetched from proceedings:
https://hot-interconnects.slack.com/archives/C0154F3UDTQ/p1786739573517959
👍 2
R
Rabindra Guha (Cerio) 1:36 PM
Are the rack size defined by NVL72? Is the scale out interface for NVLink or Ethernet?
1 reply
A
Arulselvan Madhavan 1:45 PM
I think I answered this in the call. Our modeling assumptions take rack size to be 72 XPUx. Scale-out bandwidth are ethernet/infiniband over pluggable optics
K
Kazuaki Ueda 1:36 PM
I understand that the bandwidth difference between electrical and optical scale-up connection contributes to the latency. Aside from bandwidth, are the pure signal transmission and round-trip propagation delays exactly the same for both? If they differ, could you explain whether this might impact overall latency?
1 reply
T
Thomas Graham 1:45 PM
We modeled the same propagation delays in this paper. The collective operations in prefill are more dominated by link BW than latency, which is why we made this assumption. As Arul mentioned, inference decode is more sensitive to link latency, so we have an upcoming paper that explores this more.
👍 1
L
Leila Rashidi 1:37 PM
Did scale out refer to electrical scale up only?
1 reply
T
Thomas Graham 1:47 PM
Scale-out refers to going over the scale-out network - i.e. usually using 800G or 1.6T pluggables. In the optical scale-up, we look at a larger scale-up domain (i.e. up to 1152), so you don't go over the scale-out network
P
Pankaj Mehra 1:38 PM
What will CPC do to your comparison?
2 replies
T
Taylor Groves 1:45 PM
CPC is still limited by the bandwidth you can escape off the chip shoreline.
T
Taylor Groves 1:46 PM
Whereas Passage is allowing for 3D stacking to get area-bandwidth-density. Then fibers are much more bandwidth dense than a copper wire.
✅ 1
L
Leila Rashidi 1:39 PM
You had scale out in title of your slide! What does it refer to?
P
Pankaj Mehra 1:39 PM
co-packaged copper
M
Mike Capuano 1:39 PM
So to clarify, this is NVL72 scale up (copper) + scale out (pluggable optical) to get to different racks versus a single flat scale up domain using photoconics, correct?
👍 1
2 replies
L
Leila Rashidi 1:41 PM
I think so
T
Taylor Groves 1:48 PM
Yes, identical GPU Compute and Memory, and number of GPUs in the system. Just changing the size of the scale-up network. And a second system config that increases scale-up network bandwidth per GPU.