Date, time, and room will be added once confirmed.
This workshop explains how Prefill–Decode disaggregation and Mooncake KV-cache handoff allow heterogeneous resource pools to collaborate across locations. It shares thousand-accelerator commercial deployment experience and cross-building experiments that inform bandwidth and tail-latency planning for cross-region AI infrastructure.