Date, time, and room will be added once confirmed.
This workshop presents Turbo, which offloads query multicast and partial-attention aggregation to programmable switches for context-parallel LLM inference. Online aggregation, rolling pipeline state, and load-aware aggregation trees reduce network traffic and decoding latency while preserving output quality.