from    
to    
search  

 


硬X射线谱学技术发展和应用
可再生能源科学沙龙 第一次校友论坛:高可再生能源比例下的新型电力系统
全球变化科学紫荆论坛第426期:基于气候相似性分析的台风特征及致灾研究
C3PU: An Automated Portable Platform for Organic Synthesis
报告题目:
Genome-scale Disk-based Suffix Tree Indexing
 报告人:
Mohammed J. Zaki
报告时间:
2007-06-19 15:00
报告地点:
FIT楼4区402
主办单位:
计算机系数据库实验室
  简介:

Mohammed J. Zaki is an Associate Professor of Computer Science at RPI. He received his Ph.D. degree in computer science from the University of Rochester in 1998. His research interests focus on developing novel data mining techniques for bioinformatics, performance mining, web mining, and so on. He has published over 100 papers on data mining, co-edited 11 books (including "Data Mining in Bioinformatics, Springer-London, 2005), served as guest-editor for several journals, served on the program committees of major international conferences, and co-chaired many workshops (BIOKDD, HPDM, DMKD, etc.) in data mining. He is currently an associate editor for IEEE Transactions on Knowledge and Data Engineering, action editor for Data Mining and Knowledge Discovery: An Int'l Journal, and editor for Scientific Programming, Int'l Journal of Data Warehousing and Mining, Int'l Journal of Computational Intelligence, and ACM SIGMOD Digital Symposium Collection. He received the National Science Foundation CAREER Award in 2001 and the Department of Energy Early Career Principal Investigator Award in 2002. He also received the ACM Recognition of Service Award in 2003, and an IEEE Certificate of Appreciation in 2005.

With the exponential growth of biological sequence databases, it has become critical to develop effective techniques for storing, querying, and analyzing these massive data. Suffix trees are widely used to solve many sequence-based problems, and they can be built in linear time and space, provided the resulting tree fits in main-memory. To index larger sequences, several external suffix tree algorithms have been proposed in recent years. However, they suffer from several problems such as susceptibility to data skew, non-scalability to genome-scale sequences, and non-existence of suffix links, which are crucial in various suffix tree based algorithms. In this paper, we target DNA sequences and propose a novel disk-based suffix tree algorithm called Trellis which effectively scales up to genome-scale sequences. Specifically, it can index the entire human genome using 2GB of memory, in about 4 hours and can recover all its suffix links within 2 hours. Trellis was compared to various state-of- the-art persistent disk-based suffix tree construction algorithms, and was shown to outperform the best previous methods, both in terms of indexing time and querying time.


 

今日相关信息
风力发电仿真技术进展-Samcef Wind Tur...
美国宪法的制定与成长
Economic Impact Analysis in Urban Pla...
The Formation and Growth of American ...
 
同类别相关信息
清华信息大讲堂第141讲:数据质量管理的...
清华信息大讲堂第140讲:Big Data - Se...
Integrating and Accessing Multiling...
数据研究院RONG系列论坛一:大数据与新...
大数据论坛——数据科学与技术
学术活动