from    
to    
search  

 


Symmetry restoration and quantum Mpemba effects in chaotic andlocalization sy...
Quantum Gases 2024
Stories of Fermions in an Optical Box
Contractive Unitary and Classical Shadow Tomography
报告题目:
Vocal Source and Vocal Tract Features: Their Speaker Discrimination Power and Applications in Speaker Recognition and Segmentation
 报告人:
Nengheng Zheng
The Chinese University of Hong Kong
报告时间:
2007-06-20 15:00
报告地点:
Room 1-312, FIT building,
主办单位:
Center for Speech and Language Technologies, Research Institute of Information Technology,
  简介:

Biography

Nengheng Zheng received the B.S. and M.S. degrees from Nanjing University, China, in 1997 and 2002, respectively, and the Ph.D. degree from the Chinese University of Hong Kong (CUHK), Hong Kong, China, in 2006, all in electrical and electronic engineering. He is now a Research Associate in the Department of Electronic Engineering, CUHK. Currently, he is a visiting scholar at CSLT, Research Institute of Information Technology, Tsinghua University.  His research interests include automatic speech and speaker recognition, speech enhancement and nonlinear speech analysis and modeling.

 

Abstract

Human speech is generally considered as output of the vocal tract filter system excited by the vocal source signal. State-of-the-art speaker recognition systems typically employ acoustic features which mainly carry vocal tract information (e.g. MFCC and LPCC). Although it has been believed that the source excitation signal carries rich speaker specific information, the effective information retrieving technique, as well as the successful implementation in different speaker recognition task, has been less exploited. In this talk, the work on deploying vocal source information for speaker recognition will be introduced, which was recently carried out in Digital Signal Processing and Speech Technology Laboratory, the Chinese University of Hong Kong. A set of new acoustic features, named Wavelet Octave Coefficients of Residues (WOCOR) will be firstly introduced, which aim at extracting the spectro-temporal information resided in the LP residual signal. WOCOR provides additional information to MFCC and can be used to supplement MFCC in speaker recognition. Then, the speaker discrimination power of WOCOR, in comparison with MFCC, under the conditions of linguistic content mismatch and data sparseness, will be presented. WOCOR shows superiority over MFCC under these conditions and, therefore, is particularly useful in speaker segmentation of telephone conversations, where content mismatch and data sparseness are rather critical.

今日相关信息
“近二十五年哥伦比亚建筑及建筑意义”展览...
Computational Metrology - A classific...
Platinum(II) Terdentate Complexes: Ph...
Model-Based Randomized Methods for Gl...
Modeling Stochastic Demand and Supply...
 
同类别相关信息
学术活动