Date, time, and room will be added once confirmed.
The talk discusses the challenges and optimizations in scaling up interconnect systems for supernodes, focusing on how architectural transitions and communication strategies impact performance. As model architectures shift to sparse/linear attention, reducing KV Cache pressure, tensor and sequence parallelism increase collective communication demands. The interconnect system must balance performance, cost, and ecosystem to support both training and inference efficiently.