Keynote

Omar Baldonado

Senior Director of Networking at Meta

Title: Lessons from networking Meta’s gigawatt-scale AI fleet

Abstract: Networking for AI Infrastructure has never evolved so quickly as it is now. This is especially true for Meta, who drives a gigawatt-scale fleet of AI accelerators across a wide range of model use cases and product groups; accelerator chips and systems; and networks and network hardware.

In this talk, Omar Baldonado, Senior Director of Networking at Meta, will reveal how Meta’s network team meets the aggressive time-to-market needs of AI capacity and the lessons learned from introducing new products and technologies into the fleet every quarter. Specifically, at Meta’s gigawatt scale, we treat this as a cross-network-domain problem, driven by model use cases and accelerator capabilities, that places requirements on the overall network. The team is constantly and quickly co-designing across multiple layers of the stack and across the whole AI infrastructure to (1) optimize performance, (2) ease technology and capacity transitions, and (3) provide flexibility. Because we are always introducing new accelerator hardware types, new network hardware and topologies, and new locations into the fleet, we consider all three of these goals simultaneously. While scale-up, scale-out, and scale-across networks utilize different technologies, they must be treated integrally in order to address these goals.

The talk is grounded in Meta’s experience designing, building, and operating the scale-up, scale-out, and scale-across networks for Prometheus, Meta’s first gigawatt-scale cluster, over the last 1-2 years. Omar will also cover learnings from working with both hardware and cloud vendors to leverage open standards to move faster in providing AI capacity to the rest of Meta.

Biography:

Omar Baldonado leads the groups that develop/operate Meta's global data center networks. These networks support all of Meta’s AI models and the Meta family of apps (Meta AI, Facebook, Instagram, WhatsApp, Messenger). These groups have developed some of the largest AI clusters in the world (with gigawatt-scale clusters on the way), and they continually share their work through open-source libraries (e.g., TorchComms for PyTorch, FBOSS for switches) and in communities like the Open Compute Project. Omar has been in networking since the early 1990s.