from    
to    
search  

 


清华大学材料科学与工程研究院《材料科学论坛》:High-Pressure Materials Synthes...
清华大学材料科学与工程研究院《材料科学论坛》:SiC-based Ceramic Nanocomposite...
生物和材料界面的多尺度核磁共振测量方法
The Energy Transition in Canada and Ontario
报告题目:
Emotional Speech: Its Production and Recognition
 报告人:
Sungbok Lee
University of Southern California
报告时间:
2007-07-11 14:30
报告地点:
Room 1-315, FIT building, Tsinghua University
主办单位:
信研院语音和语言技术研究中心
  简介:

Biography :

Professor Sungbok Lee was born in Korea in 1954. He received Ph. D degree in Biomedical Engineering from the University of Alabama at Birmingham in 1991. He worked as a research engineer for the Central Institute for the Deaf at Washington University during 1991 - 1997, and as a research consultant at AT&T and Lucent Bell Labs during 1998 - 2002. Now he is a research professor at University of Southern California. His main research interests are in speech production, speech signal processing and recognition including emotional speech.

Abstract:

While acoustic characteristics of emotional speech have been well documented in the literature, much less knowledge is available in the articulatory domain. Recently we have collected vocal tract data, simultaneously with speech recordings, using an electromagnetic articulograph (EMA) and a fast magnetic resonance imaging (MRI) technique while subjects produce utterances in four different acted emotions (anger, sadness, happiness and neutral). The data have been analyzed in terms of the kinematics of the tongue tip and jaw movements as a function of emotion. Articulatory timing as well as difference in the vocal tract shaping are also investigated using the functional data analysis (FDA) techniques. It is confirmed that emotion affects the tongue tip positioning and jaw opening. Furthermore the tongue tip movement range and velocity vary significantly from emotion to emotion, which implies that spectral contrasts exist among the emotions. In summary, speakers encode emotional change by modifying the place and manner of articulation while preserving linguistic identities of words. Accordingly, it is hypothesized that a robust acoustic model set trained with neutral or normal speech be effective for discriminating emotional categories. The hypothesis is tested with neutral speech HMMs trained with the TIMIT database and tested with two emotional speech databases: one is microphone and the other telephone speech. About 65% and 63% of discrimination accuracies are achieved, respectively, which supports the hypothesis. Interestingly, the plain mel-filterbank output features outperformed the mel-frequencies cepstral coefficients in that task. Detailed procedures and results will be presented and discussed.

今日相关信息
 
同类别相关信息
清华-罗姆国际产学连携论坛2018
Current Sharing for Resonant Conver...
Convex Optimization in Power System...
清华大学-华盛顿大学重点国际合作研究项...
2018清华五道口全球金融论坛
学术活动