from    
to    
search  

 


International Workshop
Phase Change-Induced Liquid Droplet Actuations on Structured Surfaces –Appli...
物理系colloquium: 阿秒脉冲的产生、应用及发展展望
天文系 Colloquium: Astrochemistry: from the interstellar medium to thecradle ...
报告题目:
Mining salient patterns in high dimensional data
 报告人:
Wei Wang
the University of North Carolina,  associate professor 
报告时间:
2007-06-29 15:00
报告地点:
FIT楼3区502
主办单位:
计算机系数据库实验室
  简介:

Wei Wang is an associate professor in the Department of Computer Science and a member of the Carolina Center for Genomic Sciences at the University of North Carolina at Chapel Hill. She received a MS degree from the State University of New York at Binghamton in 1995 and a PhD degree in Computer Science from the University of California at Los Angeles in 1999. She was a research staff member at the IBM T. J. Watson Research Center between 1999 and 2002. Dr. Wang's research interests include data mining, bioinformatics, and databases. She has filed seven patents, and has published one monograph and more than 100 research papers in international journals and major peer-reviewed conference proceedings. Dr. Wang received the IBM Invention Achievement Awards in 2000 and 2001. She was the recipient of a UNC Junior Faculty Development Award in 2003 and an NSF Faculty Early Career Development (CAREER) Award in 2005. She was named a Microsoft Research New Faculty Fellow in 2005. She was honored with the 2007 Phillip and Ruth Hettleman Prize for Artistic and Scholarly Achievement at UNC. Dr. Wang is an associate editor of the IEEE Transactions on Knowledge and Data Engineering and ACM Transactions on Knowledge Discovery in Data, and an editorial board member of the International Journal of Data Mining and Bioinformatics. She serves on the program committees of prestigious international conferences such as ACM SIGMOD, ACM SIGKDD, VLDB, ICDE, EDBT, ACM CIKM, SDM, IEEE ICDM, and SSDBM.

内容简介

The advances of new technologies have made data collection easier and faster, resulting in large and complex datasets consisting of hundreds of thousands of objects with hundreds of dimensions. Scalable and efficient unsupervised clustering methods have been the most popular approaches in analyzing these large datasets. Traditional clustering approaches typically partition objects into disjoint groups based on distances in full dimensional space. However, more often than not, some dimensions of high dimensional data may be irrelevant to a cluster and can mask the cluster’s existence. This phenomenon, called the curse of dimensionality, prevents salient structures from being discovered by traditional clustering approaches. We developed unsupervised clustering approaches to capture pattern-preserving and noise-tolerant clusters in the subspaces of high dimensional space. The proposed subspace clustering algorithms tackle the curse of dimensionality by localizing the search of clusters in the subspaces of the original high dimensional data. They go beyond the existing distance-based clustering criteria by revealing consistent patterns that can be far apart in distance.

 

 

 

今日相关信息
Digital Holography -- A New Paradigm ...
半导体纳米线、纳米器件最新研究进展(详见...
Single photon detection with avalanch...
高等研究中心学术报告:The Thermodynam...
 
同类别相关信息
清华信息大讲堂:Green Multi-Homing ...
国家实验室青年创新基金学术沙龙-网络化...
【清华五道口金融家大讲堂】Finding O...
Attacks and Defenses for Website Fi...
清华RONGv2.0系列论坛之“社会关系网络...
学术活动