报告人:张耀宇(上海交通大学)
时 间:2026年1月23日15:00
地 点:海韵园行政楼C503
内容摘要:
Condensation (also known as quantization, clustering, or alignment) is a widely observed phenomenon where neurons in the same layer tend to align with one another during the nonlinear training of deep neural networks (DNNs). It is a key characteristic of the feature learning process of neural networks. In recent years, to advance the mathematical understanding of condensation, we uncover structures regarding the dynamical regime, loss landscape and generalization for deep neural networks, based on which a novel theoretical framework emerges. This presentation will cover these findings in detail. First, I will present results regarding the dynamical regime identification of condensation at the infinite width limit, where small initialization is crucial. Then, I will discuss the mechanism of condensation at the initial training stage and the global loss landscape structure underlying condensation in later training stages, highlighting the prevalence of condensed critical points and global minimizers. Finally, I will present results on the quantification of condensation and its generalization advantage, which includes a novel estimate of sample complexity in the best-possible scenario. These results underscore the effectiveness of the phenomenological approach to understanding DNNs, paving a way for further developing deep learning theory.
个人简介:
张耀宇,上海交通大学长聘教轨副教授,2012年于上海交通大学致远学院获应用物理学学士学位及应用数学第二专业,2016年于上海交通大学数学科学学院获数学博士学位,2016年至2020年先后于纽约大学阿布扎比分校&库朗研究所、普林斯顿高等研究院从事博士后研究工作,2020年加入上海交通大学。他的研究领域聚焦深度学习的理论基础,代表性成果包括深度学习的凝聚现象与动力学态的相图分析、损失景观的嵌入原则、样本效率的乐观估计等。相关工作发表于JMLR、NeurIPS等机器学习权威期刊和会议。
联系人:陈黄鑫