from    
to    
search  

 


环境学术沙龙第659期:Anaerobic digestion of agricultural wastes: is thistechno...
理学院科学之美讲坛:探索生命系统中的设计原理
跨文化传播政治经济研究视野中的网络时代与中国故事
MAS Forum 第六期:分子科学的维度
报告题目:
Mining salient patterns in high dimensional data
 报告人:
Wei Wang
the University of North Carolina,  associate professor 
报告时间:
2007-06-29 15:00
报告地点:
FIT楼3区502
主办单位:
计算机系数据库实验室
  简介:

Wei Wang is an associate professor in the Department of Computer Science and a member of the Carolina Center for Genomic Sciences at the University of North Carolina at Chapel Hill. She received a MS degree from the State University of New York at Binghamton in 1995 and a PhD degree in Computer Science from the University of California at Los Angeles in 1999. She was a research staff member at the IBM T. J. Watson Research Center between 1999 and 2002. Dr. Wang's research interests include data mining, bioinformatics, and databases. She has filed seven patents, and has published one monograph and more than 100 research papers in international journals and major peer-reviewed conference proceedings. Dr. Wang received the IBM Invention Achievement Awards in 2000 and 2001. She was the recipient of a UNC Junior Faculty Development Award in 2003 and an NSF Faculty Early Career Development (CAREER) Award in 2005. She was named a Microsoft Research New Faculty Fellow in 2005. She was honored with the 2007 Phillip and Ruth Hettleman Prize for Artistic and Scholarly Achievement at UNC. Dr. Wang is an associate editor of the IEEE Transactions on Knowledge and Data Engineering and ACM Transactions on Knowledge Discovery in Data, and an editorial board member of the International Journal of Data Mining and Bioinformatics. She serves on the program committees of prestigious international conferences such as ACM SIGMOD, ACM SIGKDD, VLDB, ICDE, EDBT, ACM CIKM, SDM, IEEE ICDM, and SSDBM.

内容简介:

The advances of new technologies have made data collection easier and faster, resulting in large and complex datasets consisting of hundreds of thousands of objects with hundreds of dimensions. Scalable and efficient unsupervised clustering methods have been the most popular approaches in analyzing these large datasets. Traditional clustering approaches typically partition objects into disjoint groups based on distances in full dimensional space. However, more often than not, some dimensions of high dimensional data may be irrelevant to a cluster and can mask the cluster’s existence. This phenomenon, called the curse of dimensionality, prevents salient structures from being discovered by traditional clustering approaches. We developed unsupervised clustering approaches to capture pattern-preserving and noise-tolerant clusters in the subspaces of high dimensional space. The proposed subspace clustering algorithms tackle the curse of dimensionality by localizing the search of clusters in the subspaces of the original high dimensional data. They go beyond the existing distance-based clustering criteria by revealing consistent patterns that can be far apart in distance.

 

 

 

今日相关信息
Digital Holography -- A New Paradigm ...
半导体纳米线、纳米器件最新研究进展(详见...
Single photon detection with avalanch...
高等研究中心学术报告:The Thermodynam...
 
同类别相关信息
掌握沟通艺术、自信表达自我——范红教授...
首届清华信息-台大电机学科前沿技术研讨...
Statistical Model Checking for Cybe...
Data Aggregation in Communication N...
All-Pairs Shortest Paths in O(n2) E...
学术活动