Date, time, and room will be added once confirmed.
Using warehouse safety monitoring as a case study, this talk presents a hybrid architecture that runs latency-sensitive vision-language inference on edge GPUs while delegating summarization and conversational retrieval to cloud agents. It covers workload placement, model selection, optimization, shared memory, and cross-environment observability.