from    
to    
search  

 


Protein Mechanics: from Single Molecule Force Spectroscopy toProtein-based Bi...
化工系膜中心学术论坛-MOF Chemistry: From design strategies to Applications
清华大学材料科学与工程研究院《材料科学论坛》学术报告:Multi-aspect characteri...
Brain-like spiking neural networks: A 4th generation of neural network models
报告题目:
Mining salient patterns in high dimensional data
 报告人:
Wei Wang
the University of North Carolina,  associate professor 
报告时间:
2007-06-29 15:00
报告地点:
FIT楼3区502
主办单位:
计算机系数据库实验室
  简介:

Wei Wang is an associate professor in the Department of Computer Science and a member of the Carolina Center for Genomic Sciences at the University of North Carolina at Chapel Hill. She received a MS degree from the State University of New York at Binghamton in 1995 and a PhD degree in Computer Science from the University of California at Los Angeles in 1999. She was a research staff member at the IBM T. J. Watson Research Center between 1999 and 2002. Dr. Wang's research interests include data mining, bioinformatics, and databases. She has filed seven patents, and has published one monograph and more than 100 research papers in international journals and major peer-reviewed conference proceedings. Dr. Wang received the IBM Invention Achievement Awards in 2000 and 2001. She was the recipient of a UNC Junior Faculty Development Award in 2003 and an NSF Faculty Early Career Development (CAREER) Award in 2005. She was named a Microsoft Research New Faculty Fellow in 2005. She was honored with the 2007 Phillip and Ruth Hettleman Prize for Artistic and Scholarly Achievement at UNC. Dr. Wang is an associate editor of the IEEE Transactions on Knowledge and Data Engineering and ACM Transactions on Knowledge Discovery in Data, and an editorial board member of the International Journal of Data Mining and Bioinformatics. She serves on the program committees of prestigious international conferences such as ACM SIGMOD, ACM SIGKDD, VLDB, ICDE, EDBT, ACM CIKM, SDM, IEEE ICDM, and SSDBM.

内容简介

The advances of new technologies have made data collection easier and faster, resulting in large and complex datasets consisting of hundreds of thousands of objects with hundreds of dimensions. Scalable and efficient unsupervised clustering methods have been the most popular approaches in analyzing these large datasets. Traditional clustering approaches typically partition objects into disjoint groups based on distances in full dimensional space. However, more often than not, some dimensions of high dimensional data may be irrelevant to a cluster and can mask the cluster’s existence. This phenomenon, called the curse of dimensionality, prevents salient structures from being discovered by traditional clustering approaches. We developed unsupervised clustering approaches to capture pattern-preserving and noise-tolerant clusters in the subspaces of high dimensional space. The proposed subspace clustering algorithms tackle the curse of dimensionality by localizing the search of clusters in the subspaces of the original high dimensional data. They go beyond the existing distance-based clustering criteria by revealing consistent patterns that can be far apart in distance.

 

 

 

今日相关信息
Digital Holography -- A New Paradigm ...
半导体纳米线、纳米器件最新研究进展(详见...
Single photon detection with avalanch...
高等研究中心学术报告:The Thermodynam...
 
同类别相关信息
Are there insights from Quantum Co...
Correlations: From Classical to Qua...
清华信息大讲堂第149讲:Wireless Comm...
国家实验室青年创新基金学术沙龙——大数...
陈岱孙经济学纪念讲座:全球经济面临的增...
学术活动