2015/11/19 by Abhimanyu Dubey, Nikhil Naik, Dubey, Abhimanyu +7 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Image Enhancement Techniques #Machine Learning (cs.LG) #Video Surveillance and Tracking Methods #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.1511.06147
8 pages, 5 figures, In submission to IEEE TPAMI (Transactions on Pattern Analysis and Machine Intelligence)
arxiv created 2015/11/19 · openalex publication_date 2015/11/19 · arxiv updated 2015/11/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a method for learning from streaming visual data using a compact, constant size representation of all the data that was seen until a given moment. Specifically, we construct a 'coreset' representation of streaming data using a parallelized algorithm, which is an approximation of a set with relation to the squared distances between this set and all other points in its ambient space. We learn an adaptive object appearance model from the coreset tree in constant time and logarithmic space and use it for object tracking by detection. Our method obtains excellent results for object tracking on three standard datasets over more than 100 videos. The ability to summarize data efficiently makes our method ideally suited for tracking in long videos in presence of space and time constraints. We demonstrate this ability by outperforming a variety of algorithms on the TLD dataset with 2685 frames on average. This coreset based learning approach can be applied for both real-time learning of small, varied data and fast learning of big data.