vLLM Workshop

Scaling vLLM System-Wide: From Network I/O to Distributed Scheduling

Date, time, and room will be added once confirmed.

Talk overview

This workshop presents deep vLLM optimizations for cold-start latency, KV-cache efficiency, and distributed scheduling. It covers GPUDirect RDMA weight loading, stateless inference with advanced prefix caching, I/O preloading, and an event-driven no-wait scheduler for higher concurrency and GPU utilization.