KVCDN Workshop

Best Practices for Pooling LLM KV Cache with FlexKV

Date, time, and room will be added once confirmed.

Talk overview

This workshop explains how FlexKV manages and compresses KV cache across GPU memory, CPU memory, SSDs, and distributed storage. It covers cross-node pooling, asynchronous transfer, namespace isolation, and integrations with vLLM, SGLang, TensorRT-LLM, and Dynamo, drawing on practical enterprise deployment experience.