Thursday, August 13, 2026
Ryan Scherbarth (nvidia) joined the channel
Friday, August 14, 2026
Redfire joined the channel
Sunday, August 16, 2026
Gunethra joined the channel
Monday, August 17, 2026
alnonecat joined the channel
Tuesday, August 18, 2026
Steve Glaser (NVIDIA) joined the channel
Dr ABANDA EVA Pierre Robert joined the channel
David Ozog joined the channel
Allen Baum joined the channel
Taylor Groves joined the channel
Sayan Ghosh joined the channel
Pepper Marts joined the channel
Marc Cohn joined the channel
Dan Pitt joined the channel
Rohit Zambre joined the channel
Matthew Fricke joined the channel
Yiltan Temucin joined the channel
Matthew Fricke renamed the channel from "2026-d1-0915-lessons-from-networking-metas-gigawatt-scale-ai-fleet" to "2026-d1-0915-keynote-lessons-from-networking-metas-gigawatt-scale-ai-fleet"
aysebilgehan_baspinar joined the channel
Kapil Shrikhande joined the channel
Subhadeep Bhattacharya joined the channel
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-0915-keynote-lessons-from-networking-metas-gigawatt-scale-ai-fleet" to "2026-d1-0915-keynote1-lessons-from-networking-metas-gigawatt-scale-ai-fleet"
Ryan Scherbarth (nvidia) renamed the channel from "2026-d1-0915-keynote1-lessons-from-networking-metas-gigawatt-scale-ai-fleet" to "2026-d1-0915-keynote-1-lessons-from-networking-metas-gigawatt-scale-ai-fleet"
Wednesday, August 19, 2026
W
conference proceedings username and password does not work
7 replies
Y
@Ryan Scherbarth (nvidia) is there a typo in the zoom prompt
R
@Matthew Fricke / @Subhadeep Bhattacharya I'm not sure where this is configured off the top of my head
W
just followed the link and type the username/password. it does not work
S
Fixed the lobby announcement.
โ
1
W
thanks. working now. missed the "//" in password in earlier announcement
๐ 1
S
Hello Interconnects Folks
D
Weikai, please post this issue in #C0154F3UDTQ.
M
Can you be more specific on "as large as possible" for the scale up domain? Assume a time frame of 2028+, if the network and technologies could support it, would this be over 1k GPU packages? And what bandwidth per GPU package would you desire?
D
I will make sure Omar gets to these questions in the Q&A.
R
As scale-up, scale-out, and scale-across become more integrated, do you think theyโll still stay remain fundamentally different networking domains, or are we moving toward a common architecture?
C
@Omar Baldonado (Meta) If a new AL accelerator ASIC is ready to plug into Meta datacenter, what kind of network link or protocol enable the chip can plug-n-play w/o much customization at your side? Or what kind of network link do you suggest we use in such chips?
2 replies
C
Thanks a lot @Omar Baldonado (Meta)! Ethernet connection for RDMA, plus some in-tray specific local links ... like PCIe or UCIe?
O
That's a pretty good summary - first you have to see for that ASIC what type of server/rack system it is sitting in (and what traffic requirements the ASIC will have in/out) - basically what scale-up is it going to sit in (or if it sits in one at all, vs. being part of say a regular front-end fabric). Then you can see what requirements need to go in/out of the ASIC and that informs the building-level fabric requirements (whether a backend scale-out fabric or a front-end fabric).
C
@Omar Baldonado (Meta) Does Meta's AI specific links for AI can distinguish the traffic is for memory or computing? AI chips require different traffic character for different challenges. Thank you for sharing your insights!
1 reply
O
"For memory" - do you mean like what is RDMA traffic vs non-RDMA traffic? If so, then yes, we have telemetry from all the switches that is monitoring the traffic (protocol/endpoints/...)
S
Aside from power delivery distance, how much does distance/accessibility to natural resources like water source matter in determining the breadth and potential expansion of the scale across domain?
C
HI @Dan Pitt So nice to see you again! โค๏ธ It's amazing HotI has made huge impacts to hardware industry every year! Thank you and HOTI committee for your exceptional leadership! ๐ ๐
๐ 2
P
With scale-in and vendor proprietary integration growing to the scale of proprietary racks, are you not putting pressure on Ethernet to do things it was not designed to do?
1 reply
O
The number of historical 802.x standards should we've continually been putting pressure on Ethernet over the decades ๐ But seriously, not all of Ethernet is appropriate for scale-up/low-latency/high-bandwidth efficiency, and I wouldn't put an Ethernet switch in the middle of a compute tray for die-to-die communication if I didn't have to. We push for Ethernet because it provides a well-known, mature ecosystem of products, vendors, and researchers, even as we all push it to do new things.
๐ 1
M
Do you have an example of upgrading (or a plan to upgrade) the scale-out backend network with newer switches? What lessons were learned or how are you thinking about this?
1 reply
O
Meta had a paper at SIGCOMM a few years ago (2023?) that gave examples of how we plan and do "drains" of fabrics. A lot of that is applied to our scale-out network. Something we also push for is warm-boot and ISSU-style capabilities - you can see within the OCP Switch Abstraction Interface (SAI) project of how we ask for this, because it lets us do these upgrades with less impact.
H
Can we have one standard for scale up instead of having NVLINK, UALINK, ESUN and SUE.
1 reply
O
Trying ๐ One of the takeaways I shared is that scale-up is heavily part of the AI system and that varies heavily. ESUN is a multi-vendor initiative within OCP, including a number of those AI vendors. Small note: I believe SUE has moved to be more focused on transport layer (SUE-T), even as it leverages Ethernet underneath.
A
Really useful point that new accelerators, network hardware and topologies are entering the fleet so quickly. When one layer changes, say GPU/NIC, PCIe or CXL, transport, collective stack, or topology, how do you determine the smallest set of cross-layer assumptions that must be revalidated before rollout? Is there a machine-readable dependency model for that today, or is it still largely expert-driven? Thanks for sharing!
3 replies
O
Great question - internally, we've continually automate/standardize this as much as possible, with higher workload-level benchmarking/test suites as well as component-level benchmarks/test suites. Also, note there is a tradeoff between time to bring up capacity and time to optimize it. So in certain situations, we may take X% of roofline performance and have it working sooner, and then spend time optimizing it in parallel (rather than delaying the capacity, waiting for 100% roofline). This is because often there is so much to learn operationally when inserting a new component, that I want the experience with it as soon as possible.
A
That makes sense, Omar, getting safe capacity online early can be worth more than waiting for the last few percent. I am wondering then what defines the "safe to run real jobs" gate, and do issues found during bring-up normally become regression tests for future rollouts?
O
We have a few different levels of qualification gates defined, so that folks that want/get "early access" know that we're going to be working some things out still . And those issues indeed get folded in for future verification test suites for future capacity,
P
You haven't talked about ROUTING. Can you comment on differences versus traditional networking? Thanks
3 replies
O
I love routing (back in early 2000's, I worked for a company named RouteScience!). In scale-up, the routing problem is pretty easy ๐ and in scale-across, there's definitely a similar routing problem. We use BGP actively even in our scale-out fabrics, but where we've had to extend the routing control plane some is to make sure failure handling is fast and propagated well. There's also a lot of work in load balancing (ECMP) to further ensure failures given all the parallel links we have.
L
@Omar Baldonado (Meta) , does Meta uses MRC?
B
What is the most critical hardware component prone to fail first in a networking fleet? Is is the Laser, GPU/CPU chip, photonics chip, or the board components?
1 reply
O
It's based on some older generation infrastructure from 2023/2024, but we shared some of AI infrastructure reliability learnings in our lengthy Llama3 paper (
arxiv.org/pdf/2407.21783). Section 3.3 covers some of the failure rates/components.
S
What is a Time-expanded Network (TEN), and why is it central to the TACOS framework's ability to synthesize topology-aware collective algorithms?
1 reply
O
Without going into the specifics of this framework, I can say that it relates to some of the lessons I presented in the talk. For example, the bubble elimination in a collective by overlapping comms and compute is an example where cross-layer/whole system knowledge is required. Topology-awareness is critical for jobs/collectives, especially when dealing with scale-across networks where there is much more heterogeneity in the network (as opposed to scale-up and scale-out). There are multiple ways to bring this sort of awareness into the researchers' development workflow.
S
How are SD-WAN vendors adapting to the requirements of scale across HPC networks?
1 reply
O
I'm not totally aware of this - I know how the switch/network ASIC vendors are responding, but we have our own in-house SD-WAN controller system.
C
How do you secure programmability in the scale out and scale up networks as protocols like MRC, RoCEv2 evolve rapidly? Have you thought about FPGA-based NICs, to not be reliant on ASIC development.
2 replies
O
Meta will be talking more about some of our thoughts on protocols and NICs in our Networking @ Scale conference next week.
atscaleconference.com/events/networking-2026 - encourage you all to attend! Regarding FPGA-based NICs, yes we have looked at them, but those come with their pros/cons. Often in these technologies, we'll evaluate and use some at a certain scale, even as we push for more scalable/performant solutions in parallel.
R
What is the expected timeline for CPO adoption in scale-up architectures, beyond its initial use in scale-out network switches?
1 reply
O
As you probably saw and refer to, we've been very active in evaluating CPO in general as a technology (initially in scale-out - where we announced some reliability results at ECOC last year). We also announced earlier the OCI MSA, which is related. Optics in scale-up has a clear value prop assuming you are moving to a certain size scale-up (which we want).
C
If you have any experience with heterogenous/disaggregated inference compute/deployments, where you combine GPU/XPUs from different vendors, how do you bridge the different scale-up and scale-out protocols?
2 replies
O
I think heterogeneity in the fleet (at different layers like switches, NICs, etc) is different from GPU/XPU interoperability. I think that is a really hard problem and not just because of network but the whole software stack, numerics accuracy & compatibility, etc. In a way, the network protocols are an easier part of the problem.
L
Does Meta uses disaggregated inference?
C
Thanks a lot @Omar Baldonado (Meta) for sharing your valuable insights! :rose:
โ
1
R
You mentioned that upgrading to newer generations of hardware is a challenge. In practice, do you usually phase in the new systems while older ones keep running, or do power, cooling, and network changes sometimes force you to rebuild larger sections at once?
2 replies
O
It really depends on how interoperable the previous generation of hardware is. upgrading to the next-gen switch or NIC can keep leverage a lot of previous infrastructure, but later AI-system-level upgrades aren't compatible or cost effective with the existing building (e.g., we do liquid cooling in buildings without facility-level cooling, but we have to tradeoff cost/benefit/feasibility of such upgrades)
R
I see that makes sense Omar thank you so much for the reply.
R
Has META looked at AI workload performance based upon various server with variable number of GPUs, within a server, to determine which, Scale Up vs Scale Out, is more beneficial.
G
great presentation. How broadly might we see Scale Across deployed? Are we talking about Training applications only (therefore likely Scale Across would require handfuls of DCs connected)... or do you expect it would be used for Inference (which might imply many more Scale Across locations connected)?
1 reply
O
A lot of plans are focused on training - indeed, our Llama4 pretraining was done using the scale-across connecting buildings of our 129K cluster and we continue to do that. However, "inference" has a lot of different parts now between GPUs, CPUs, and storage. Hard right now to say how broad that will go.
S
also, please use #C0154F3UDTQ channel
O
HI y'all - thanks for attending and engaging with so many questions! As you can probably tell, I'm excited about all of this networking we have ahead of us!
O
I'm going through questions now
Y
Hi Omar, thanks for the great talk! Is recording of your talk available for replay later? thank you
D
All videos will be available on the platform in a few days. Then the entire conference will be viewable on YouTube indefinitely. You may check out the HotI YouTube channel to see past talks.
๐ 1
O
Thanks again to all of you and to @Dan Pitt and all the HotInterconnects coordinators- see you all at the rest of the conference!
๐ 3๐ 1
Friday, August 21, 2026
G
@Omar Baldonado (Meta) very insightful keynote, Thanks. Would you, or a colleague, be interested in a short chat about Teradio's wireless OOB management network? 100G, Ethernet compatible, and no wires, truly out-of-band. We're a spin-out from Northeastern U. and have US Gov $$$ to develop an adaptive wireless infrastructure for AI Data Centers. Our patented sub-THz pencil-beam technology makes it all work. It's actually an interesting architectural enhancement as well for other uses. Love to chat. Best. Greg Whelan CEO Teradio