from    
to    
search  

 


Protein Mechanics: from Single Molecule Force Spectroscopy toProtein-based Bi...
化工系膜中心学术论坛-MOF Chemistry: From design strategies to Applications
清华大学材料科学与工程研究院《材料科学论坛》学术报告:Multi-aspect characteri...
Brain-like spiking neural networks: A 4th generation of neural network models
报告题目:
Emotional Speech: Its Production and Recognition
 报告人:
Sungbok Lee
University of Southern California
报告时间:
2007-07-11 14:30
报告地点:
Room 1-315, FIT building, Tsinghua University
主办单位:
信研院语音和语言技术研究中心
  简介:

Biography :

Professor Sungbok Lee was born in Korea in 1954. He received Ph. D degree in Biomedical Engineering from the University of Alabama at Birmingham in 1991. He worked as a research engineer for the Central Institute for the Deaf at Washington University during 1991 - 1997, and as a research consultant at AT&T and Lucent Bell Labs during 1998 - 2002. Now he is a research professor at University of Southern California. His main research interests are in speech production, speech signal processing and recognition including emotional speech.

Abstract:

While acoustic characteristics of emotional speech have been well documented in the literature, much less knowledge is available in the articulatory domain. Recently we have collected vocal tract data, simultaneously with speech recordings, using an electromagnetic articulograph (EMA) and a fast magnetic resonance imaging (MRI) technique while subjects produce utterances in four different acted emotions (anger, sadness, happiness and neutral). The data have been analyzed in terms of the kinematics of the tongue tip and jaw movements as a function of emotion. Articulatory timing as well as difference in the vocal tract shaping are also investigated using the functional data analysis (FDA) techniques. It is confirmed that emotion affects the tongue tip positioning and jaw opening. Furthermore the tongue tip movement range and velocity vary significantly from emotion to emotion, which implies that spectral contrasts exist among the emotions. In summary, speakers encode emotional change by modifying the place and manner of articulation while preserving linguistic identities of words. Accordingly, it is hypothesized that a robust acoustic model set trained with neutral or normal speech be effective for discriminating emotional categories. The hypothesis is tested with neutral speech HMMs trained with the TIMIT database and tested with two emotional speech databases: one is microphone and the other telephone speech. About 65% and 63% of discrimination accuracies are achieved, respectively, which supports the hypothesis. Interestingly, the plain mel-filterbank output features outperformed the mel-frequencies cepstral coefficients in that task. Detailed procedures and results will be presented and discussed.

今日相关信息
 
同类别相关信息
清华信息大讲堂169讲:Recent Advance...
清华论坛第70讲:新型城镇化的挑战与机...
国际雾计算产学研联盟——中国(北京)雾...
清华大学校企合作委员会2017海外年会
清华信息大讲堂170讲:从失败样本学习...
学术活动