vLLM Workshop

From Local Acceleration to System-Wide Optimization: Full-Stack LLM Inference in Practice

Date, time, and room will be added once confirmed.

Talk overview

Drawing on Hunyuan inference and GPU engineering experience, this workshop presents a service-goal-oriented method for LLM optimization. It moves from bottleneck analysis and kernel optimization to coordinated compute, memory, communication, and scheduling, emphasizing the tradeoffs required to turn local speedups into measurable end-to-end gains.