1. The most prevalent form of data in today’s digital world are images and videos. A most important goal of computer vision in artificial intelligence research is to develop algorithms that can recognize hundreds of thousands of objects in millions and billions of images. To achieve this goal, researchers need to work with millions and billions of images annotated with accurate labels for training and benchmarking different algorithms. Up till only a few years ago, limited by labor and money resource, computer vision scientists had been working with datasets consisted of only thousands of images annotated across a few dozen object classes. In 2008, our lab started a project called ImageNet, aiming to build the largest annotated image dataset in computer vision research. Our goal was to put together a dataset of tens of millions of images annotated with object classes found in the English dictionary (about twenty thousand of them!), an impossible mission using traditional ways of hiring subjects in university campuses. Instead we used a crowdsourcing technology to recruit tens of thousands of online workers to help us labeling more than half billion images using the Amazon Mechanical Turk platform. This talk will give a detailed account of our experience of this exciting and emerging technology and what we have done to build the largest image dataset in our research community. Furthermore, we will highlight a number of research projects that take advantage of the large-scale ImageNet data, including benchmarking today’s computer vision algorithms, as well as a new large-scale image recognition system that is built to never make mistakes.
2. The web is not perfect: while text is easily searched and organized, pictures (the vast majority of the bits that one can find on-line) are not. In order to see how one could improve the web and make pictures first-class citizens of the web, I explore the idea of Visipedia, a visual interface for Wikipedia that is able to answer visual queries andenables experts to contribute and organize visual knowledge. Four distinct groups of humans would interact through Visipedia: users, experts, visual workers and machine vision scientists. The latter would gradually build automata able to interpret images. I will explore some of the technical challenges involved in making Visipedia happen and present our initial results in crowdsourcing visual annotation, building automated field guides that may be deployed on mobile devices and combining machines and humans for discovering, harvesting and organizingvisual information.