from    
to    
search  

 


Classification and construction of crystalline topologicalsuperconductors an...
韩国影视节目“走出去”的历程: 从平台与内容的角度看国际传播
未来音乐将会怎样?
清华大学材料科学与工程研究院《材料科学论坛》:Adaptive Nanophotonics by Novel...
报告题目:
清华软件论坛第17期|林宙辰-Adan: Adaptive Nesterov Momentum Algorithm for FasterOptimizing Deep Models
 报告人:
林宙辰
北京大学智能学院教授
报告时间:
2023-05-11 10:00
报告地点:
东主楼10区316室,腾讯会议号:151-197-186
主办单位:
410#软件学院
  简介:

Different kinds of deep networks typically need different optimizers, which have to be chosen after ?multiple trials, making the training process inefficient. To relieve ?this issue and consistently improve the model training speed across deep networks, we propose the ADAptive Nesterov momentum algorithm ?(Adan). Adan first reformulates the vanilla Nesterov acceleration to develop a new Nesterov momentum estimation (NME) method, then adopts ?NME to estimate the first- and second-order moments of the gradient ?for convergence acceleration. Besides, we prove that Adan finds an ?-approximate first-order stationary point within $O(?^{?3.5})$ stochastic gradient complexity on the non-convex stochastic problems, matching the best-known lower bound. Extensive experimental results show that Adan consistently surpasses the corresponding SoTA optimizers on vision, language, and RL tasks and sets new SoTAs for ?many popular networks, e.g. ResNet, ConvNext, ViT, Swin, MAE, DETR, ?GPT-2, Transformer-XL. More surprisingly, Adan can use half of the ?epochs of SoTA optimizers to achieve higher or comparable performance on ViT, GPT-2, MAE, etc, and also shows great tolerance to a large range of minibatch size, e.g. from 1k to 32k. Code is released at https://github.com/sail-sg/Adan.

今日相关信息
集成电路系列学术邀请报告——Chiplet设...
Phase Separation in Synapse Formation...
天文系Colloquium:Physical Mechanisms...
【数学之美-杰出学者讲坛】2023年第4期 ...
物理系colloquium: 低维量子材料的超快电...
 
同类别相关信息
清华软件论坛第十五期|程鸿: Solving ...
【千帆讲堂-青年医师科学家学术沙龙】医...
AIR学术工作坊第3期|大模型时代AI生物...
清华软件论坛第十四期 | 戎珂:数据要...
清华软件论坛 | 区块链的发展形势与区...
学术活动