Thursday, August 13, 2026
Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026
Arulselvan Madhavan joined the channel
Redfire joined the channel
Sunday, August 16, 2026
Gunethra joined the channel
Monday, August 17, 2026
alnonecat joined the channel
Tuesday, August 18, 2026
Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
David Ozog joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Dan Pitt joined the channel
Matthew Fricke joined the channel
Yiltan Temucin joined the channel
Rohit Zambre joined the channel
aysebilgehan_baspinar joined the channel
Kapil Shrikhande joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-1040-network-design-at-scale" to "2026-d1-1040-paper-network-design-at-scale"
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-1040-paper-network-design-at-scale" to "2026-d1-1040-paper-session-a-network-design-at-scale"
Wednesday, August 19, 2026
K
Welcome to the session : Network Design and Scale.
K
Feel free to enter your questions here.
M
Do you see MRLS as applicable to back-end scale out networks? It seems like increasing diameter beyond 2 would not be desirable and that MRLS would only be applicable to the front end network.
R
This sound similar to Rockport Networks Torus Networking technology.
1 reply
L
Is this posted in website of Cerio?
P
Does this use stndard routing protocols or is there a new routing protocol?
4 replies
R
I think any packet based protocol, that has a SRC and DST address. Thus Ethernet is a supported protocol.
👍 1
P
he mentioned you need to do 2 lookup's that new?
A
The underlying minimal routing tables can be built using standard mechanisms, but what changes is how packets are forwarded. Usually, standard routings do 1 lookup to the routing table. Polarized routing needs 2 lookups of the minimal routing table: one querying the source and the other for the destination. It requires both to compute the correct ports.
K
Does your cost model take into account differences in copper vs. optical interconnect. And how does the MRLS topology compare to Fat-tree in its breakdown of optical vs. copper links .
S
Is it possible to do parallel simulation in CAMINOS, because with more endpoints, serial simulation overhead can be significant? Also, have you considered including the topology/routing on existing discrete event simulators like SST/Macro (which may support parallel simulation)?
L
What is hardware requirement in switch?
C
@Alejandro Cano So MRLS expands the network to new nodes with the trade off to reduce the bandwidth of existing nodes? Is it possible to add another layer of fat tree at top when the workload is too much, so it's 2D MRLS?
L
Does it run over commodity switches?
L
What about per pkt state?
K
Please enter your questions for the second Paper of the session.
L
Any insight regarding the impact on power for this study with optical?
3 replies
T
4.3pJ/b used based of hardware from the lab. That's all in SerDes, laser, optics.
T
You'd compare that against your long-reach 224G SerDes from your favorite source for the electrical link. If it's a pluggable optics its going to be much higher ~16pJ/b.
👍 1
L
Great, thank for the details
C
Do you think some of the assumptions in the analysis will change with the introduction of XPO formfactore, increasing the density of pluggable optics for a more classic system (copper/optics)
1 reply
T
XPO is a good density solution keep pluggables relevant for a while longer, but not going to be as efficient as CPO.
K
Is there a reason you showed the scale-up domain up to 1152 GPUs ? Where is the constraint coming from – is it from the optics, o# racks, distance ?
1 reply
M
It is limited by switch radix
👍 1
N
Can someone share the link to the paper?
B
Do you need more fault-tolerant designs for photonic interconnects?
1 reply
T
You do add more components (e.g. laser + PIC) with photonics, but there are lots of options for making these reliable and replaceable. External laser, detatchable fibers, redundancy in design, etc.
C
The papers are up in the conference proceedings
R
Are the rack size defined by NVL72? Is the scale out interface for NVLink or Ethernet?
1 reply
A
I think I answered this in the call. Our modeling assumptions take rack size to be 72 XPUx. Scale-out bandwidth are ethernet/infiniband over pluggable optics
K
I understand that the bandwidth difference between electrical and optical scale-up connection contributes to the latency. Aside from bandwidth, are the pure signal transmission and round-trip propagation delays exactly the same for both? If they differ, could you explain whether this might impact overall latency?
1 reply
T
We modeled the same propagation delays in this paper. The collective operations in prefill are more dominated by link BW than latency, which is why we made this assumption. As Arul mentioned, inference decode is more sensitive to link latency, so we have an upcoming paper that explores this more.
👍 1
L
Did scale out refer to electrical scale up only?
1 reply
T
Scale-out refers to going over the scale-out network - i.e. usually using 800G or 1.6T pluggables. In the optical scale-up, we look at a larger scale-up domain (i.e. up to 1152), so you don't go over the scale-out network
P
What will CPC do to your comparison?
2 replies
T
CPC is still limited by the bandwidth you can escape off the chip shoreline.
T
Whereas Passage is allowing for 3D stacking to get area-bandwidth-density. Then fibers are much more bandwidth dense than a copper wire.
✅ 1
L
You had scale out in title of your slide! What does it refer to?
M
So to clarify, this is NVL72 scale up (copper) + scale out (pluggable optical) to get to different racks versus a single flat scale up domain using photoconics, correct?
👍 1
2 replies
T
Yes, identical GPU Compute and Memory, and number of GPUs in the system. Just changing the size of the scale-up network. And a second system config that increases scale-up network bandwidth per GPU.