简介: |
报告摘要:With the advancements of next-generation sequencing technology, it is now possible to study samples directly obtained from the environment. Particularly, 16S rRNA gene sequences have been frequently used to profile the diversity of organisms underlying an environmental sample. However, such studies are still taxed to determine both the number of operational taxonomic units (OTUs) and their relative abundance in a sample. To address these challenges, we propose an unsupervised Bayesian clustering method termed Clustering 16S rRNA for OTU Prediction (CROP). CROP can find clusters based on the natural organization of data without setting a hard cutoff threshold (3%/5%) as required by hierarchical clustering methods. By applying our method to several datasets, we demonstrate that CROP is robust against sequencing errors and that it produces more accurate results than conventional hierarchical clustering methods. The second generation sequencing technology also enables the study of the whole transcriptome by shotgun sequencing of RNA sequences. In practice, the de novo transcriptome assembly is the only choice to the study of the transcriptome of organisms that do not have reference genome sequences, and it can also be applied to identify novel transcripts and structural variations in the gene regions of model organisms. The de novo transcriptome assembly differs from the de novo genome assembly in that it assembles multiple alternatively spliced gene transcripts simultaneously instead of individual chromosome sequences. We propose a new de novo transcriptome assembly called WEAV which first partitions RNA-seq reads into clusters, and then for each cluster, applies a de Bruijn graph based method to simultaneously identify multiple alternatively spliced gene transcripts.
报告人简介:Professor Ting Chen received his Ph.D. in Computer Science from Computer Science Department, SUNY at Stony Brook in 1997 and is now an associate professor of Computer Science in University of Southern California,. Within the fields of computational biology and bioinformatics, Professor Chen seeks to apply computer algorithms and mathematical methods to answer questions in biology and medicine, specifically in studies of human genetics, proteomics, and genomics. Dr. Chen's research includes (1) large-scale sequence analysis for the next generation sequencing data, (2) analysis of genetic variations of the human genome for their relationships to human diseases, (3) studies of functions and dynamic of large-scale biochemical networks inside the cell, and (4) identification of proteins through analyzing mass spectrometry data. Detailed information about Dr. Chen can be found at http://www.cmb.usc.edu/people/tingchen/
|