from    
to    
search  

 


天文系 Colloquium: On the Cosmic Baryon Cycle---Insights Learned from theCosm...
可再生能源科学沙龙第二期:在中美气候合作背景下解析美国能源部科技创新战略
全球变化科学紫荆论坛第434期:事件遥感新视角
车辆与运载学院299期学术沙龙-从时空轨迹到自动驾驶的频谱建模:构建一个“十级”框...
报告题目:
Emotional Speech: Its Production and Recognition
 报告人:
Sungbok Lee
University of Southern California
报告时间:
2007-07-11 14:30
报告地点:
Room 1-315, FIT building, Tsinghua University
主办单位:
信研院语音和语言技术研究中心
  简介:

Biography :

Professor Sungbok Lee was born in Korea in 1954. He received Ph. D degree in Biomedical Engineering from the University of Alabama at Birmingham in 1991. He worked as a research engineer for the Central Institute for the Deaf at Washington University during 1991 - 1997, and as a research consultant at AT&T and Lucent Bell Labs during 1998 - 2002. Now he is a research professor at University of Southern California. His main research interests are in speech production, speech signal processing and recognition including emotional speech.

Abstract:

While acoustic characteristics of emotional speech have been well documented in the literature, much less knowledge is available in the articulatory domain. Recently we have collected vocal tract data, simultaneously with speech recordings, using an electromagnetic articulograph (EMA) and a fast magnetic resonance imaging (MRI) technique while subjects produce utterances in four different acted emotions (anger, sadness, happiness and neutral). The data have been analyzed in terms of the kinematics of the tongue tip and jaw movements as a function of emotion. Articulatory timing as well as difference in the vocal tract shaping are also investigated using the functional data analysis (FDA) techniques. It is confirmed that emotion affects the tongue tip positioning and jaw opening. Furthermore the tongue tip movement range and velocity vary significantly from emotion to emotion, which implies that spectral contrasts exist among the emotions. In summary, speakers encode emotional change by modifying the place and manner of articulation while preserving linguistic identities of words. Accordingly, it is hypothesized that a robust acoustic model set trained with neutral or normal speech be effective for discriminating emotional categories. The hypothesis is tested with neutral speech HMMs trained with the TIMIT database and tested with two emotional speech databases: one is microphone and the other telephone speech. About 65% and 63% of discrimination accuracies are achieved, respectively, which supports the hypothesis. Interestingly, the plain mel-filterbank output features outperformed the mel-frequencies cepstral coefficients in that task. Detailed procedures and results will be presented and discussed.

今日相关信息
 
同类别相关信息
大海捞毒针——网络安全高级持续威胁(A...
Blockchain vis-à-vis Database Sys...
An Introduction to AI and Deep Lear...
Intelligent Software Engineering: S...
弹性云计算平台性能和成本的优化
学术活动