We are at the dawn the golden age of statistical machine learning and big data, with new powerful methods proliferating and an ever-increasing array of useful applications. However, a key bottleneck to applying machine learning to many practical problems is the paucity of accurately-labeled training data and the cost in time and money to obtain human expert judgments or to conduct definitive scientific experiments. Active learning strives to find the most informative data instances to label by an external “oracle”, but to achieve widespread practicality we must do more, i.e.: cope with multiple external information sources (experts, crowds, experiments, observations, etc.), estimate their reliability (accuracy) and their availability and their cost, and jointly optimize selection of data instances and sources of expertise in an amortized setting to maximize learning in any given time horizon. This joint-optimization, especially necessary in big data analytics, is called Proactive Learning; we discuss how to do it for increasingly complex scenario. Then we touch on several applications, including crowd-source learning for machine translation, and proteomic host-pathogen analysis.
讲演者简介
Dr. Jaime Carbonell is the Director of the Language Technologies Institute and Allen Newell Professor of Computer Science at Carnegie Mellon University. He received SB degrees in Physics and Mathematics from MIT, and MS and PhD degrees in Computer Science from Yale University. His current research includes machine learning, scalable data mining, text mining, machine translation and computational proteomics. He invented Proactive Machine Learning, including its underlying decision-theoretic framework. He is also known for the Maximal Marginal Relevance principle in information retrieval, for derivational analogy in problem solving and for example-based machine translation and for machine learning in structural biology, and in protein interaction networks. Overall, he has published some 330 papers and books and supervised over 50 PhD dissertations. Dr. Carbonell has served on multiple governmental advisory committees such as the Human Genome Committee of the National Institutes of Health, the Oakridge National Laboratories Scientific Advisory Board, the National Institute of Standards and Technology Interactive Systems Scientific Advisory Board, and the German National Artificial Intelligence (DFKI) Scientific Advisory Board. In education, Carbonell created the PhD and MS degrees in Language Technologies at CMU and designed courses in language technologies, machine learning, data sciences and electronic commerce.