from    
to    
search  

 


Symmetry restoration and quantum Mpemba effects in chaotic andlocalization sy...
Quantum Gases 2024
Stories of Fermions in an Optical Box
Contractive Unitary and Classical Shadow Tomography
报告题目:
清华软件论坛第17期|林宙辰-Adan: Adaptive Nesterov Momentum Algorithm for FasterOptimizing Deep Models
 报告人:
林宙辰
北京大学智能学院教授
报告时间:
2023-05-11 10:00
报告地点:
东主楼10区316室,腾讯会议号:151-197-186
主办单位:
410#软件学院
  简介:

Different kinds of deep networks typically need different optimizers, which have to be chosen after ?multiple trials, making the training process inefficient. To relieve ?this issue and consistently improve the model training speed across deep networks, we propose the ADAptive Nesterov momentum algorithm ?(Adan). Adan first reformulates the vanilla Nesterov acceleration to develop a new Nesterov momentum estimation (NME) method, then adopts ?NME to estimate the first- and second-order moments of the gradient ?for convergence acceleration. Besides, we prove that Adan finds an ?-approximate first-order stationary point within $O(?^{?3.5})$ stochastic gradient complexity on the non-convex stochastic problems, matching the best-known lower bound. Extensive experimental results show that Adan consistently surpasses the corresponding SoTA optimizers on vision, language, and RL tasks and sets new SoTAs for ?many popular networks, e.g. ResNet, ConvNext, ViT, Swin, MAE, DETR, ?GPT-2, Transformer-XL. More surprisingly, Adan can use half of the ?epochs of SoTA optimizers to achieve higher or comparable performance on ViT, GPT-2, MAE, etc, and also shows great tolerance to a large range of minibatch size, e.g. from 1k to 32k. Code is released at https://github.com/sail-sg/Adan.

今日相关信息
集成电路系列学术邀请报告——Chiplet设...
Phase Separation in Synapse Formation...
天文系Colloquium:Physical Mechanisms...
【数学之美-杰出学者讲坛】2023年第4期 ...
物理系colloquium: 低维量子材料的超快电...
 
同类别相关信息
人工智能拓展火灾安全研究的进展
第四届清华信息前沿交叉论坛
浅谈人工智能重塑城市公共安全治理新范式
AIR学术沙龙第37期|创新智能环境:无...
脑机接口时代,我们还能做什么?——脑科...
学术活动