Date, time, and room will be added once confirmed.
The talk examines intrinsic security risks in large language models and explains why output-level constraints cannot guarantee trustworthy behavior. It introduces the theory of equivalent interaction mechanism interpretation and its applications to formal mechanism auditing, intrinsic security hardening, and training-efficiency optimization. The session also covers findings on risky mechanisms in open-source models, model generalization, and mechanism-alignment training.