Agentic AI on Edge

KTransformers: The Fine-Tuning-to-Inference Closed Loop for Full-Precision 1T-Class MoE Models on Consumer Hardware

Date, time, and room will be added once confirmed.

Talk overview

KTransformers enables efficient local inference for 1T-class MoE models on consumer hardware by leveraging CPU-GPU heterogeneous computing. It supports a full fine-tuning-to-inference loop, allowing model serving, feedback collection, and deployment back to the system. This approach makes large-scale model iteration accessible to small teams and researchers without high hardware costs.