Date, time, and room will be added once confirmed.
Drawing on Hunyuan inference and GPU engineering experience, this workshop presents a service-goal-oriented method for LLM optimization. It moves from bottleneck analysis and kernel optimization to coordinated compute, memory, communication, and scheduling, emphasizing the tradeoffs required to turn local speedups into measurable end-to-end gains.
Speakers