from    
to    
search  

 


清华大学材料科学与工程研究院《材料科学论坛》:近红外二区磷光成像
清华大学材料科学与工程研究院《材料科学论坛》:四方Sr4Al2O7:一种用于制备高质量...
Massive Particles at Spatial Infinity
An update on holography of information (Review)
报告题目:
清华软件论坛第17期|林宙辰-Adan: Adaptive Nesterov Momentum Algorithm for FasterOptimizing Deep Models
 报告人:
林宙辰
北京大学智能学院教授
报告时间:
2023-05-11 10:00
报告地点:
东主楼10区316室,腾讯会议号:151-197-186
主办单位:
410#软件学院
  简介:

Different kinds of deep networks typically need different optimizers, which have to be chosen after ?multiple trials, making the training process inefficient. To relieve ?this issue and consistently improve the model training speed across deep networks, we propose the ADAptive Nesterov momentum algorithm ?(Adan). Adan first reformulates the vanilla Nesterov acceleration to develop a new Nesterov momentum estimation (NME) method, then adopts ?NME to estimate the first- and second-order moments of the gradient ?for convergence acceleration. Besides, we prove that Adan finds an ?-approximate first-order stationary point within $O(?^{?3.5})$ stochastic gradient complexity on the non-convex stochastic problems, matching the best-known lower bound. Extensive experimental results show that Adan consistently surpasses the corresponding SoTA optimizers on vision, language, and RL tasks and sets new SoTAs for ?many popular networks, e.g. ResNet, ConvNext, ViT, Swin, MAE, DETR, ?GPT-2, Transformer-XL. More surprisingly, Adan can use half of the ?epochs of SoTA optimizers to achieve higher or comparable performance on ViT, GPT-2, MAE, etc, and also shows great tolerance to a large range of minibatch size, e.g. from 1k to 32k. Code is released at https://github.com/sail-sg/Adan.

今日相关信息
集成电路系列学术邀请报告——Chiplet设...
Phase Separation in Synapse Formation...
天文系Colloquium:Physical Mechanisms...
【数学之美-杰出学者讲坛】2023年第4期 ...
物理系colloquium: 低维量子材料的超快电...
 
同类别相关信息
中性团簇红外光谱研究
生成式人工智能的技术演进与行业应用
清华软件论坛第20期|俞士纶(Philip S...
清华软件论坛第十九期|C.Mohan: Syste...
AIR学术工作坊|AI赋能智能机器人理解...
学术活动