Tuesday, August 18, 2026
Ryan Scherbarth (nvidia) joined the channel
David Ozog joined the channel
Sayan Ghosh joined the channel
Jiaqi Lou joined the channel
bsahu joined the channel
Fabian joined the channel
Francois Labonte joined the channel
Daud Arslan joined the channel
Shahbozbek Hakimov joined the channel
Tong Xu joined the channel
myoung-gyun.suh joined the channel
xi.chen joined the channel
Mahfuz joined the channel
Varun Chotalia joined the channel
pierre-louis.benard joined the channel
Matthew Fricke joined the channel
Advm joined the channel
ahmadrazarehman033 joined the channel
Howard Wang joined the channel
rashid-ahmed.kukkady joined the channel
Wednesday, August 19, 2026
L
How CPUs are used with MTIA for agentic AI usecases?
1 reply
K
Unfortunately I can't share details here because this is part of currently active discussions.
M
Are those RDMA NICs infiniband or ethernet? If ethernet RoCE or something else. .
3 replies
K
Yes, these are RDMA / RoCE NICs.
L
Has Meta any plan to use UFH rather than ESUN 1.0 header?
H
what is difference between MTIA and GPU
A
Could you provide some examples where this fungibility for scale-up vs scale-out helps?
3 replies
K
As mentioned in the slides, we could potentially use all the NICs for scale-up connectivity depending on the workload requirements.
👍 1
A
Thanks for the response, do you have any insights on workload requirements in terms of scale-up/scale-out ratio?
K
Generally speaking training requires collectives over accelerators in multiple racks and inference jobs tend to be smaller in scale resulting in different requirements.
M
What SerDes are you using?
1 reply
K
I'm unable to share this information at this time.
R
Is the choice between Message Engine offload and direct PE injection made dynamically based on message size or latency, or is it decided ahead of time in software?
3 replies
L
You may find your response in MITA paper accepted for super computing 2026. Arxiv version is available
R
Thank you Leila for the response
K
I'd also recommend
arxiv.org/html/…. As for the question - typically its a decision made up front on the host side. However, dynamic decision in device is also possible and something we are exploring.
L
Is it possible to converge scale up and out in future? Does Meta experience reduction of scale up bandwidth to scale out bandwidth ratio? 5 is much lower than 10, which has been observed in industry in the past
L
Is ESUN header used for scale up? Is there any header optimization for scale up? Are you running RoCE over ESUN?