Date, time, and room will be added once confirmed.
This talk reviews GLM's collaboration with inference communities including SGLang and vLLM. It covers release-day model support, speculative decoding, quantization, Prefill–Decode disaggregation, accuracy alignment, performance tuning, and how upstream contributions and open technical reports make high-quality inference practices available to the broader ecosystem.