2026-d1-0915-keynote-1-lessons-from-networking-metas-gigawatt-scale-ai-fleet

Archived conversation · Aug 13, 2026 8:38 PM – Aug 21, 2026 8:03 AM · 91 messages
Thursday, August 13, 2026

Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026

Redfire joined the channel
Sunday, August 16, 2026

Gunethra joined the channel
Monday, August 17, 2026

alnonecat joined the channel
Tuesday, August 18, 2026

Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
David Ozog joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Dan Pitt joined the channel
Rohit Zambre joined the channel
Matthew Fricke joined the channel
Yiltan Temucin joined the channel
Matthew Fricke renamed the channel from "2026-d1-0915-lessons-from-networking-metas-gigawatt-scale-ai-fleet" to "2026-d1-0915-keynote-lessons-from-networking-metas-gigawatt-scale-ai-fleet"
aysebilgehan_baspinar joined the channel
Kapil Shrikhande joined the channel
Subhadeep Bhattacharya joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-0915-keynote-lessons-from-networking-metas-gigawatt-scale-ai-fleet" to "2026-d1-0915-keynote1-lessons-from-networking-metas-gigawatt-scale-ai-fleet"
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-0915-keynote1-lessons-from-networking-metas-gigawatt-scale-ai-fleet" to "2026-d1-0915-keynote-1-lessons-from-networking-metas-gigawatt-scale-ai-fleet"
Wednesday, August 19, 2026

O
Omar Baldonado (Meta) 10:49 AM
Hi everyone!
A
Ahmed Khalil (unaffiliated) 10:57 AM
Hello all!
R
Ramtin Soleymani 11:00 AM
Hello!
W
Weikai Sun 11:06 AM
conference proceedings username and password does not work
W
Weikai Sun 11:06 AM
image.png
7 replies
Y
Yiltan Temucin 11:08 AM
conf26// should work
Y
Yiltan Temucin 11:09 AM
@Ryan Scherbarth (nvidia) is there a typo in the zoom prompt
R
Ryan Scherbarth (nvidia) 11:10 AM
@Matthew Fricke / @Subhadeep Bhattacharya I'm not sure where this is configured off the top of my head
W
Weikai Sun 11:11 AM
just followed the link and type the username/password. it does not work
Y
Yiltan Temucin 11:12 AM
@Weikai Sun Just to double check, does the prompt match the note in the #C0154F3UDTQ channel?

https://hot-interconnects.slack.com/archives/C0154F3UDTQ/p1786739573517959
S
Subhadeep Bhattacharya 11:13 AM
Fixed the lobby announcement.
โœ… 1
W
Weikai Sun 11:13 AM
thanks. working now. missed the "//" in password in earlier announcement
๐Ÿ‘ 1
S
saurabh(photonic researcher) 11:08 AM
Hello Interconnects Folks
D
Dan Pitt 11:24 AM
Weikai, please post this issue in #C0154F3UDTQ.
M
Mike Capuano 11:30 AM
Can you be more specific on "as large as possible" for the scale up domain? Assume a time frame of 2028+, if the network and technologies could support it, would this be over 1k GPU packages? And what bandwidth per GPU package would you desire?
D
Dan Pitt 11:41 AM
I will make sure Omar gets to these questions in the Q&A.
R
Ram Sharan Chaulagain 11:47 AM
As scale-up, scale-out, and scale-across become more integrated, do you think theyโ€™ll still stay remain fundamentally different networking domains, or are we moving toward a common architecture?
C
Cissy Yuan 11:50 AM
@Omar Baldonado (Meta) If a new AL accelerator ASIC is ready to plug into Meta datacenter, what kind of network link or protocol enable the chip can plug-n-play w/o much customization at your side? Or what kind of network link do you suggest we use in such chips?
2 replies
C
Cissy Yuan 12:11 PM
Thanks a lot @Omar Baldonado (Meta)! Ethernet connection for RDMA, plus some in-tray specific local links ... like PCIe or UCIe?
O
Omar Baldonado (Meta) 12:40 PM
That's a pretty good summary - first you have to see for that ASIC what type of server/rack system it is sitting in (and what traffic requirements the ASIC will have in/out) - basically what scale-up is it going to sit in (or if it sits in one at all, vs. being part of say a regular front-end fabric). Then you can see what requirements need to go in/out of the ASIC and that informs the building-level fabric requirements (whether a backend scale-out fabric or a front-end fabric).
D
Dan Pitt 11:51 AM
Hi Cissy!
C
Cissy Yuan 11:52 AM
@Omar Baldonado (Meta) Does Meta's AI specific links for AI can distinguish the traffic is for memory or computing? AI chips require different traffic character for different challenges. Thank you for sharing your insights!
1 reply
O
Omar Baldonado (Meta) 12:42 PM
"For memory" - do you mean like what is RDMA traffic vs non-RDMA traffic? If so, then yes, we have telemetry from all the switches that is monitoring the traffic (protocol/endpoints/...)
S
Sayan Ghosh 11:54 AM
Aside from power delivery distance, how much does distance/accessibility to natural resources like water source matter in determining the breadth and potential expansion of the scale across domain?
C
Cissy Yuan 11:55 AM
HI @Dan Pitt So nice to see you again! โค๏ธ It's amazing HotI has made huge impacts to hardware industry every year! Thank you and HOTI committee for your exceptional leadership! ๐Ÿ™Œ ๐Ÿ‘
๐Ÿ‘ 2
P
Pankaj Mehra 11:59 AM
With scale-in and vendor proprietary integration growing to the scale of proprietary racks, are you not putting pressure on Ethernet to do things it was not designed to do?
1 reply
O
Omar Baldonado (Meta) 12:45 PM
The number of historical 802.x standards should we've continually been putting pressure on Ethernet over the decades ๐Ÿ™‚ But seriously, not all of Ethernet is appropriate for scale-up/low-latency/high-bandwidth efficiency, and I wouldn't put an Ethernet switch in the middle of a compute tray for die-to-die communication if I didn't have to. We push for Ethernet because it provides a well-known, mature ecosystem of products, vendors, and researchers, even as we all push it to do new things.
๐Ÿ‘ 1
M
Mike Capuano 11:59 AM
Do you have an example of upgrading (or a plan to upgrade) the scale-out backend network with newer switches? What lessons were learned or how are you thinking about this?
1 reply
O
Omar Baldonado (Meta) 12:50 PM
Meta had a paper at SIGCOMM a few years ago (2023?) that gave examples of how we plan and do "drains" of fabrics. A lot of that is applied to our scale-out network. Something we also push for is warm-boot and ISSU-style capabilities - you can see within the OCP Switch Abstraction Interface (SAI) project of how we ask for this, because it lets us do these upgrades with less impact.
H
Hesham ElBakoury 12:00 PM
Can we have one standard for scale up instead of having NVLINK, UALINK, ESUN and SUE.
1 reply
O
Omar Baldonado (Meta) 12:52 PM
Trying ๐Ÿ™‚ One of the takeaways I shared is that scale-up is heavily part of the AI system and that varies heavily. ESUN is a multi-vendor initiative within OCP, including a number of those AI vendors. Small note: I believe SUE has moved to be more focused on transport layer (SUE-T), even as it leverages Ethernet underneath.
A
Ahmed Khalil (unaffiliated) 12:04 PM
Really useful point that new accelerators, network hardware and topologies are entering the fleet so quickly. When one layer changes, say GPU/NIC, PCIe or CXL, transport, collective stack, or topology, how do you determine the smallest set of cross-layer assumptions that must be revalidated before rollout? Is there a machine-readable dependency model for that today, or is it still largely expert-driven? Thanks for sharing!
3 replies
O
Omar Baldonado (Meta) 12:56 PM
Great question - internally, we've continually automate/standardize this as much as possible, with higher workload-level benchmarking/test suites as well as component-level benchmarks/test suites. Also, note there is a tradeoff between time to bring up capacity and time to optimize it. So in certain situations, we may take X% of roofline performance and have it working sooner, and then spend time optimizing it in parallel (rather than delaying the capacity, waiting for 100% roofline). This is because often there is so much to learn operationally when inserting a new component, that I want the experience with it as soon as possible.
A
Ahmed Khalil (unaffiliated) 4:21 AM
That makes sense, Omar, getting safe capacity online early can be worth more than waiting for the last few percent. I am wondering then what defines the "safe to run real jobs" gate, and do issues found during bring-up normally become regression tests for future rollouts?
O
Omar Baldonado (Meta) 2:21 PM
We have a few different levels of qualification gates defined, so that folks that want/get "early access" know that we're going to be working some things out still . And those issues indeed get folded in for future verification test suites for future capacity,
P
Pankaj Mehra 12:07 PM
You haven't talked about ROUTING. Can you comment on differences versus traditional networking? Thanks
3 replies
O
Omar Baldonado (Meta) 1:06 PM
I love routing (back in early 2000's, I worked for a company named RouteScience!). In scale-up, the routing problem is pretty easy ๐Ÿ™‚ and in scale-across, there's definitely a similar routing problem. We use BGP actively even in our scale-out fabrics, but where we've had to extend the routing control plane some is to make sure failure handling is fast and propagated well. There's also a lot of work in load balancing (ECMP) to further ensure failures given all the parallel links we have.
L
Leila Rashidi 2:29 PM
@Omar Baldonado (Meta) , does Meta uses MRC?
O
Omar Baldonado (Meta) 2:16 PM
I'd encourage you to come to our networking @scale event next week to see what we're doing in this transport area! atscaleconference.com/events/networking-2026
๐Ÿ‘ 1
B
bsahu 12:07 PM
What is the most critical hardware component prone to fail first in a networking fleet? Is is the Laser, GPU/CPU chip, photonics chip, or the board components?
1 reply
O
Omar Baldonado (Meta) 1:18 PM
It's based on some older generation infrastructure from 2023/2024, but we shared some of AI infrastructure reliability learnings in our lengthy Llama3 paper (arxiv.org/pdf/2407.21783). Section 3.3 covers some of the failure rates/components.
S
syedmuavizurrehman019 12:08 PM
What is a Time-expanded Network (TEN), and why is it central to the TACOS framework's ability to synthesize topology-aware collective algorithms?
1 reply
O
Omar Baldonado (Meta) 1:23 PM
Without going into the specifics of this framework, I can say that it relates to some of the lessons I presented in the talk. For example, the bubble elimination in a collective by overlapping comms and compute is an example where cross-layer/whole system knowledge is required. Topology-awareness is critical for jobs/collectives, especially when dealing with scale-across networks where there is much more heterogeneity in the network (as opposed to scale-up and scale-out). There are multiple ways to bring this sort of awareness into the researchers' development workflow.
S
sankar ramamoorthi 12:09 PM
How are SD-WAN vendors adapting to the requirements of scale across HPC networks?
1 reply
O
Omar Baldonado (Meta) 1:25 PM
I'm not totally aware of this - I know how the switch/network ASIC vendors are responding, but we have our own in-house SD-WAN controller system.
C
Christoffer Wang Bjรธrnsen 12:11 PM
How do you secure programmability in the scale out and scale up networks as protocols like MRC, RoCEv2 evolve rapidly? Have you thought about FPGA-based NICs, to not be reliant on ASIC development.
2 replies
O
Omar Baldonado (Meta) 1:28 PM
Meta will be talking more about some of our thoughts on protocols and NICs in our Networking @ Scale conference next week. atscaleconference.com/events/networking-2026 - encourage you all to attend! Regarding FPGA-based NICs, yes we have looked at them, but those come with their pros/cons. Often in these technologies, we'll evaluate and use some at a certain scale, even as we push for more scalable/performant solutions in parallel.
C
Christoffer Wang Bjรธrnsen 1:36 PM
Thank you
R
Rayun Kim 12:11 PM
What is the expected timeline for CPO adoption in scale-up architectures, beyond its initial use in scale-out network switches?
1 reply
O
Omar Baldonado (Meta) 1:35 PM
As you probably saw and refer to, we've been very active in evaluating CPO in general as a technology (initially in scale-out - where we announced some reliability results at ECOC last year). We also announced earlier the OCI MSA, which is related. Optics in scale-up has a clear value prop assuming you are moving to a certain size scale-up (which we want).
C
Christoffer Wang Bjรธrnsen 12:15 PM
If you have any experience with heterogenous/disaggregated inference compute/deployments, where you combine GPU/XPUs from different vendors, how do you bridge the different scale-up and scale-out protocols?
2 replies
O
Omar Baldonado (Meta) 1:39 PM
I think heterogeneity in the fleet (at different layers like switches, NICs, etc) is different from GPU/XPU interoperability. I think that is a really hard problem and not just because of network but the whole software stack, numerics accuracy & compatibility, etc. In a way, the network protocols are an easier part of the problem.
L
Leila Rashidi 2:35 PM
Does Meta uses disaggregated inference?
C
Cissy Yuan 12:16 PM
Thanks a lot @Omar Baldonado (Meta) for sharing your valuable insights! :rose:
โœ… 1
R
Ramtin Soleymani 12:17 PM
You mentioned that upgrading to newer generations of hardware is a challenge. In practice, do you usually phase in the new systems while older ones keep running, or do power, cooling, and network changes sometimes force you to rebuild larger sections at once?
2 replies
O
Omar Baldonado (Meta) 1:43 PM
It really depends on how interoperable the previous generation of hardware is. upgrading to the next-gen switch or NIC can keep leverage a lot of previous infrastructure, but later AI-system-level upgrades aren't compatible or cost effective with the existing building (e.g., we do liquid cooling in buildings without facility-level cooling, but we have to tradeoff cost/benefit/feasibility of such upgrades)
R
Ramtin Soleymani 4:46 PM
I see that makes sense Omar thank you so much for the reply.
R
Rabindra Guha (Cerio) 12:19 PM
Has META looked at AI workload performance based upon various server with variable number of GPUs, within a server, to determine which, Scale Up vs Scale Out, is more beneficial.
G
gnotter 12:22 PM
great presentation. How broadly might we see Scale Across deployed? Are we talking about Training applications only (therefore likely Scale Across would require handfuls of DCs connected)... or do you expect it would be used for Inference (which might imply many more Scale Across locations connected)?
1 reply
O
Omar Baldonado (Meta) 1:51 PM
A lot of plans are focused on training - indeed, our Llama4 pretraining was done using the scale-across connecting buildings of our 129K cluster and we continue to do that. However, "inference" has a lot of different parts now between GPUs, CPUs, and storage. Hard right now to say how broad that will go.
S
syedmuavizurrehman019 12:26 PM
NEXT SESSION TIME?
S
Sayan Ghosh 12:29 PM
see hoti.org/2026/program.html (10:40A PST after break)
S
Sayan Ghosh 12:29 PM
also, please use #C0154F3UDTQ channel
O
Omar Baldonado (Meta) 12:37 PM
HI y'all - thanks for attending and engaging with so many questions! As you can probably tell, I'm excited about all of this networking we have ahead of us!
O
Omar Baldonado (Meta) 12:38 PM
I'm going through questions now
Y
ysun 12:42 PM
Hi Omar, thanks for the great talk! Is recording of your talk available for replay later? thank you
D
Dan Pitt 12:46 PM
All videos will be available on the platform in a few days. Then the entire conference will be viewable on YouTube indefinitely. You may check out the HotI YouTube channel to see past talks.
๐Ÿ™ 1
O
Omar Baldonado (Meta) 1:52 PM
Thanks again to all of you and to @Dan Pitt and all the HotInterconnects coordinators- see you all at the rest of the conference!
๐Ÿ™Œ 3๐Ÿ‘ 1
Friday, August 21, 2026

G
gwhelan 8:03 AM
@Omar Baldonado (Meta) very insightful keynote, Thanks. Would you, or a colleague, be interested in a short chat about Teradio's wireless OOB management network? 100G, Ethernet compatible, and no wires, truly out-of-band. We're a spin-out from Northeastern U. and have US Gov $$$ to develop an adaptive wireless infrastructure for AI Data Centers. Our patented sub-THz pencil-beam technology makes it all work. It's actually an interesting architectural enhancement as well for other uses. Love to chat. Best. Greg Whelan CEO Teradio