Date, time, and room will be added once confirmed.
This workshop examines how agent workloads transform KV-cache reuse, retention, and capacity requirements. It introduces UCM, a lifecycle-aware cache system spanning GPU memory, DRAM, SSD, and distributed storage, with adaptive policies, persistent KV storage, tier-aware placement, and optimized data movement for lower inference cost.