from    
to    
search  

 


Symmetry restoration and quantum Mpemba effects in chaotic andlocalization sy...
Quantum Gases 2024
Stories of Fermions in an Optical Box
Contractive Unitary and Classical Shadow Tomography
报告题目:
Recent Innovations in Speech Technology Ignited by Deep Learning
 报告人:
Li Deng
Microsoft Research,Professor
报告时间:
2012-12-14 16:00
报告地点:
FIT楼,1-415
主办单位:
信研院
  简介:

Semantic information embedded in the speech signal manifests itself in a dynamic process rooted in the deep linguistic hierarchy as an intrinsic part of the human cognitive system. Modeling both the dynamic process and the deep structure for advancing speech technology has been an active pursuit for over more than 20 years, but it is only within past two to three years that technological breakthrough has been created by a methodology commonly referred to as "deep learning". Deep Belief Net (DBN) and the related deep neural nets are recently being used to supersede the Gaussian mixture model component in HMM-based speech recognition, and has produced dramatic error rate reduction in both phone recognition and large vocabulary speech recognition of industry scale while keeping the HMM component intact. On the other hand, the (constrained) Dynamic Bayesian Networks have been developed for many years to improve the dynamic models of speech aimed to overcome the IID assumption as a key weakness of the HMM, with a set of techniques commonly known as hidden dynamic/trajectory models or articulatory-like segmental representations. A history of these two largely separate lines of research will be critically reviewed and analyzed in the context of modeling the deep and dynamic linguistic hierarchy for advancing speech recognition technology.

The first wave of innovation has successfully unseated Gaussian mixture model and MFCC-like features --- two of the three main pillars of the 20-year-old technology in speech recognition. Future directions will be discussed and analyzed on supplanting the final pillar --- HMM --- where frame-level scores are to be enhanced to dynamic-segment scores through new waves of innovation capitalizing on multiple lines of research that has enriched our knowledge of the deep, dynamic process of human speech.

今日相关信息
云计算及软件工业的商业模型(高水平英文课...
载药磁性纳米复合介质用于肿瘤磁感应热化疗
清华-爱泼斯坦国际传播系列讲座(二)En...
环境学术沙龙第111期:水环境污染生物预警...
 
同类别相关信息
人工智能拓展火灾安全研究的进展
第四届清华信息前沿交叉论坛
浅谈人工智能重塑城市公共安全治理新范式
AIR学术沙龙第37期|创新智能环境:无...
有机-无机杂化二维MXene材料
学术活动