from    
to    
search  

 


金属增材制造——新兴复杂合金及多尺度结构设计 Additive Manufacturing ofEmergin...
Next generation power devices: challenges, possibilities, andopportunities in...
学堂班系列讲座:“化学反应的量子特性”
环境学术沙龙第658期:细菌耐药新机制解析
报告题目:
Emotional Speech: Its Production and Recognition
 报告人:
Sungbok Lee
University of Southern California
报告时间:
2007-07-11 14:30
报告地点:
Room 1-315, FIT building, Tsinghua University
主办单位:
信研院语音和语言技术研究中心
  简介:

Biography :

Professor Sungbok Lee was born in Korea in 1954. He received Ph. D degree in Biomedical Engineering from the University of Alabama at Birmingham in 1991. He worked as a research engineer for the Central Institute for the Deaf at Washington University during 1991 - 1997, and as a research consultant at AT&T and Lucent Bell Labs during 1998 - 2002. Now he is a research professor at University of Southern California. His main research interests are in speech production, speech signal processing and recognition including emotional speech.

Abstract:

While acoustic characteristics of emotional speech have been well documented in the literature, much less knowledge is available in the articulatory domain. Recently we have collected vocal tract data, simultaneously with speech recordings, using an electromagnetic articulograph (EMA) and a fast magnetic resonance imaging (MRI) technique while subjects produce utterances in four different acted emotions (anger, sadness, happiness and neutral). The data have been analyzed in terms of the kinematics of the tongue tip and jaw movements as a function of emotion. Articulatory timing as well as difference in the vocal tract shaping are also investigated using the functional data analysis (FDA) techniques. It is confirmed that emotion affects the tongue tip positioning and jaw opening. Furthermore the tongue tip movement range and velocity vary significantly from emotion to emotion, which implies that spectral contrasts exist among the emotions. In summary, speakers encode emotional change by modifying the place and manner of articulation while preserving linguistic identities of words. Accordingly, it is hypothesized that a robust acoustic model set trained with neutral or normal speech be effective for discriminating emotional categories. The hypothesis is tested with neutral speech HMMs trained with the TIMIT database and tested with two emotional speech databases: one is microphone and the other telephone speech. About 65% and 63% of discrimination accuracies are achieved, respectively, which supports the hypothesis. Interestingly, the plain mel-filterbank output features outperformed the mel-frequencies cepstral coefficients in that task. Detailed procedures and results will be presented and discussed.

今日相关信息
 
同类别相关信息
清华大学林家翘讲座
清华大学林家翘讲座
新世纪的台湾电影
Low Rank Approximation and Regressi...
周光召基金会获奖者清华论坛
学术活动