from    
to    
search  

 


清华化工论坛第86讲-Understanding and Controlling Zeolite Synthesis
Cooperative late transition metal pincer complexes for manufacturing B-Nbased...
碳中和背景下林木纤丝的可持续化开发
清华大学化工系绿电化工研究中心 学术论坛:Microenvironment tuning for CO2 elec...
报告题目:
New Directions in Robust Automatic Speech Recognition
 报告人:
Richard M. Stern
Prof. Carnegie Mellon University
报告时间:
2008-09-29 09:00
报告地点:
Room 1-415, FIT building, Tsinghua University
主办单位:
信研院
  简介:

2008年清华大学信息技术研究院系列学术报告13

Biography

Richard M. Stern received the S.B. degree from the Massachusetts Institute of Technology in 1970, the M.S. from the University of California, Berkeley, in 1972, and the Ph.D. from MIT in 1977, all in electrical engineering. He has been on the faculty of Carnegie Mellon University since 1977, where he is currently a Professor in the Electrical and Computer Engineering, Computer Science, and Biomedical Engineering Departments, and the Language Technologies Institute. Much of Dr. Stern's current research is in spoken language systems, where he is particularly concerned with the development of techniques with which automatic speech recognition can be made more robust with respect to changes in environment and acoustical ambience. He has also developed sentence parsing and speaker adaptation algorithms for earlier CMU speech systems. In addition to his work in speech recognition, Dr. Stern also maintains an active research program in psychoacoustics, where he is best known for theoretical work in binaural perception. Dr. Stern is a fellow of the Acoustical Society of America, the 2008-2009 Distinguished Lecturer of the International Speech Communication Association, a recipient of the Allen Newell Award for Research Excellence in 1992, and he served as General Chair of Interspeech 2006. He is also a member of the IEEE and the Audio Engineering Society. Abstract

As speech recognition technology is transferred from the laboratory to the marketplace, robustness in recognition is becoming increasingly important. This talk will review and discuss several classical and contemporary approaches to robust speech recognition. The most tractable types of environmental degradation are produced by quasi-stationary additive noise and quasi-stationary linear filtering. These distortions can be largely ameliorated by the "classical" techniques of cepstral high-pass filtering (as exemplified by cepstral mean normalization and RASTA filtering), as well as by techniques that develop statistical models of the distortion (such as codeword-dependent cepstral normalization and vector Taylor series expansion). Nevertheless, these types of approaches fail to provide much useful improvement when speech is degraded by transient or non-stationary noise such as background music or speech. We describe and compare the effectiveness toward providing increased recognition accuracy in difficult acoustical environments of techniques based on missing-feature compensation, multi-band analysis, combination of complementary streams of information, physiologically-motivated auditory processing, and auditory scene analysis.

今日相关信息
 
同类别相关信息
配体保护的金纳米团簇的结构和性质的计算...
金属氮杂环卡宾催化的资源分子高值化反应
亚纳米和纳米尺度上二维膜的制备和传输
清华2023高分子前沿讲座--有趣的带电高...
清美讲坛:乔治·史威兹贝尔的动画艺术世...
学术活动