from    
to    
search  

 


Symmetry restoration and quantum Mpemba effects in chaotic andlocalization sy...
Quantum Gases 2024
Stories of Fermions in an Optical Box
Contractive Unitary and Classical Shadow Tomography
报告题目:
清华信息大讲堂第96讲-NEC第3讲:1.Scaling Up Object Recognition: What have we done, and where are we going? 2.Visipedia: collaborative organization of visual knowledge
 报告人:
1.Fei-Fei Li 2.Pietro Perona,
1. Associate Professor,Stanford University
2. Professor,California Institute of Technology
报告时间:
2012-12-17 09:00
报告地点:
信息楼(FIT)多功能厅
主办单位:
信息学院
  简介:
1. The most prevalent form of data in today’s digital world are images and videos. A most important goal of computer vision in artificial intelligence research is to develop algorithms that can recognize hundreds of thousands of objects in millions and billions of images. To achieve this goal, researchers need to work with millions and billions of images annotated with accurate labels for training and benchmarking different algorithms. Up till only a few years ago, limited by labor and money resource, computer vision scientists had been working with datasets consisted of only thousands of images annotated across a few dozen object classes. In 2008, our lab started a project called ImageNet, aiming to build the largest annotated image dataset in computer vision research. Our goal was to put together a dataset of tens of millions of images annotated with object classes found in the English dictionary (about twenty thousand of them!), an impossible mission using traditional ways of hiring subjects in university campuses. Instead we used a crowdsourcing technology to recruit tens of thousands of online workers to help us labeling more than half billion images using the Amazon Mechanical Turk platform. This talk will give a detailed account of our experience of this exciting and emerging technology and what we have done to build the largest image dataset in our research community. Furthermore, we will highlight a number of research projects that take advantage of the large-scale ImageNet data, including benchmarking today’s computer vision algorithms, as well as a new large-scale image recognition system that is built to never make mistakes.
2. The web is not perfect: while text is easily searched and organized, pictures (the vast majority of the bits that one can find on-line) are not. In order to see how one could improve the web and make pictures first-class citizens of the web, I explore the idea of Visipedia, a visual interface for Wikipedia that is able to answer visual queries andenables experts to contribute and organize visual knowledge. Four distinct groups of humans would interact through Visipedia: users, experts, visual workers and machine vision scientists. The latter would gradually build automata able to interpret images. I will explore some of the technical challenges involved in making Visipedia happen and present our initial results in crowdsourcing visual annotation, building automated field guides that may be deployed on mobile devices and combining machines and humans for discovering, harvesting and organizingvisual information.

 

今日相关信息
Single-molecule Biophysical View of G...
清华大学新人文讲座系列之(十二)文化传承...
Uncertainty, Flexibility, Valuation &...
STM Investigation of Coulomb Impuriti...
图书馆自助服务介绍
 
同类别相关信息
人工智能拓展火灾安全研究的进展
第四届清华信息前沿交叉论坛
浅谈人工智能重塑城市公共安全治理新范式
AIR学术沙龙第37期|创新智能环境:无...
有机-无机杂化二维MXene材料
学术活动