简介: |
Advances in personalized genomic medicine have the potential to change the way we treat diseases, but our ability to translate these advances into a reality for patients will stem from our ability to successfully and rapidly analyze patient data. Next-generation sequencing (NGS) technologies are now increasingly used to find disease genes in human genomic studies. Integrating functional characterization of identified mutations with comprehensive genome interpretation could provide compelling evidence implicating new disease-contributing mutations and genes in phenotypically well-characterized patients. The rise of big data in NGS will contribute to better treatment paradigms, leading to improvements in diagnosis and targeted medications, which may ultimately lead to an overall cost-savings in health care. In spite of this, enormous challenges for analyzing large-scale NGS still exist including data storage, processing, scaling, quality control management, and interpretation. Genome sequencing of humans has been a leading contributor to Big Data, exceeding the abilities of currently used approaches to store, manage, share, analyze, and interpret it effectively. To infer biological insights from massive amounts of NGS data in a short period of time, we developed a software framework called SeqHBase to quickly identify disease-contributing genes. SeqHBase is a big data toolset for analyzing family-based sequencing data to detect de novo, inherited homozygous, or compound heterozygous mutations that may contribute to disease manifestations. It works through a distributed and parallel manner over multiple data nodes. With 20 data nodes, SeqHBase took about 5 seconds to analyze whole-exome sequencing (WES) data for a family quartet and approximately 1 minute to analyze whole-genome sequencing (WGS) data for a 10-member three-generation family. These results demonstrated SeqHBase’s high efficiency and scalability.
Dr. Min He
Dr. Min He is a tenure-track faculty member at the rank of Assistant Professor in the Center for Human Genetics at Marshfield Clinic Research Foundation (MCRF). Dr. He also holds a joint appointment in the Biomedical Informatics Research Center at MCRF and an adjunct Professorship in the Computation and Informatics in Biology and Medicine at University of Wisconsin-Madison. Prior to assuming his present positions, Dr. He served as Assistant Professor in the Center for Human Genome Variation at Duke University with responsibilities including leading development of a number of statistical methods for analyzing over 4,000 whole genome/exome sequencing data. The current focus of Dr. He’s research is on developing statistical approaches and computational tools for genomic medicine. |