2021/03/24 by Songtao He, He, Songtao, Favyen Bastani +12
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Annotation #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #FOS: Computer and information sciences #Global Positioning System #Image (mathematics) #Machine Learning (cs.LG) #Minimum bounding box #Object (grammar) #Object detection #Pattern recognition (psychology) #Pixel #TRACE (psycholinguistics) #Video Surveillance and Tracking Methods #Video tracking #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2103.13428
published in arXiv (Cornell University) (Cornell University) · https://people.csail.mit.edu/songtao/tagme.html
arxiv created 2021/03/24 · openalex publication_date 2021/03/24 · arxiv updated 2021/03/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Training high-accuracy object detection models requires large and diverse annotated datasets. However, creating these data-sets is time-consuming and expensive since it relies on human annotators. We design, implement, and evaluate TagMe, a new approach for automatic object annotation in videos that uses GPS data. When the GPS trace of an object is available, TagMe matches the object's motion from GPS trace and the pixels' motions in the video to find the pixels belonging to the object in the video and creates the bounding box annotations of the object. TagMe works using passive data collection and can continuously generate new object annotations from outdoor video streams without any human annotators. We evaluate TagMe on a dataset of 100 video clips. We show TagMe can produce high-quality object annotations in a fully-automatic and low-cost way. Compared with the traditional human-in-the-loop solution, TagMe can produce the same amount of annotations at a much lower cost, e.g., up to 110x.