2017/03/26 by Hossein Hosseini, Hosseini, Hossein, Baicen Xiao +3 · 1 voice
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.1703.09793
openalex publication_date 2017/03/26 · arxiv published 2017/03/26 · arxiv updated 2017/03/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Despite the rapid progress of the techniques for image classification, video\nannotation has remained a challenging task. Automated video annotation would be\na breakthrough technology, enabling users to search within the videos.\nRecently, Google introduced the Cloud Video Intelligence API for video\nanalysis. As per the website, the system can be used to "separate signal from\nnoise, by retrieving relevant information at the video, shot or per frame"\nlevel. A demonstration website has been also launched, which allows anyone to\nselect a video for annotation. The API then detects the video labels (objects\nwithin the video) as well as shot labels (description of the video events over\ntime). In this paper, we examine the usability of the Google's Cloud Video\nIntelligence API in adversarial environments. In particular, we investigate\nwhether an adversary can subtly manipulate a video in such a way that the API\nwill return only the adversary-desired labels. For this, we select an image,\nwhich is different from the video content, and insert it, periodically and at a\nvery low rate, into the video. We found that if we insert one image every two\nseconds, the API is deceived into annotating the video as if it only contained\nthe inserted image. Note that the modification to the video is hardly\nnoticeable as, for instance, for a typical frame rate of 25, we insert only one\nimage per 50 video frames. We also found that, by inserting one image per\nsecond, all the shot labels returned by the API are related to the inserted\nimage. We perform the experiments on the sample videos provided by the API\ndemonstration website and show that our attack is successful with different\nvideos and images.\n