from    
to    
search  

 


环境学术沙龙第704期:Advancing Separation Technologies for a Circular Battery ...
经济变革的全球策略 | “清华论坛”第106讲 暨“人文与社会”系列讲座总第110期
AIR学术沙龙第37期|创新智能环境:无线通讯和感知的新视角
【数学之美-杰出学者讲坛】2024年第1期 || What is curvature? And why is it impo...
报告题目:
清华软件论坛第17期|林宙辰-Adan: Adaptive Nesterov Momentum Algorithm for FasterOptimizing Deep Models
 报告人:
林宙辰
北京大学智能学院教授
报告时间:
2023-05-11 10:00
报告地点:
东主楼10区316室,腾讯会议号:151-197-186
主办单位:
410#软件学院
  简介:

Different kinds of deep networks typically need different optimizers, which have to be chosen after ?multiple trials, making the training process inefficient. To relieve ?this issue and consistently improve the model training speed across deep networks, we propose the ADAptive Nesterov momentum algorithm ?(Adan). Adan first reformulates the vanilla Nesterov acceleration to develop a new Nesterov momentum estimation (NME) method, then adopts ?NME to estimate the first- and second-order moments of the gradient ?for convergence acceleration. Besides, we prove that Adan finds an ?-approximate first-order stationary point within $O(?^{?3.5})$ stochastic gradient complexity on the non-convex stochastic problems, matching the best-known lower bound. Extensive experimental results show that Adan consistently surpasses the corresponding SoTA optimizers on vision, language, and RL tasks and sets new SoTAs for ?many popular networks, e.g. ResNet, ConvNext, ViT, Swin, MAE, DETR, ?GPT-2, Transformer-XL. More surprisingly, Adan can use half of the ?epochs of SoTA optimizers to achieve higher or comparable performance on ViT, GPT-2, MAE, etc, and also shows great tolerance to a large range of minibatch size, e.g. from 1k to 32k. Code is released at https://github.com/sail-sg/Adan.

今日相关信息
集成电路系列学术邀请报告——Chiplet设...
Phase Separation in Synapse Formation...
天文系Colloquium:Physical Mechanisms...
【数学之美-杰出学者讲坛】2023年第4期 ...
物理系colloquium: 低维量子材料的超快电...
 
同类别相关信息
How AI changes human-computer inter...
名师微沙龙:文科PI系列第一场
2019中国金融科技学术年会
北京信息科学与技术国家研究中心青年创新...
气候变化大讲堂第14讲:能源低碳转型与...
学术活动