2023/07/04 by Rajesh Vasa, Simmons, Anj, Vasa, Rajesh · 1 citation
Computer Science · Medicine · #Multimodal Machine Learning Applications #Viral Infections and Outbreaks Research #Anomaly Detection Techniques and Applications
paper · pdf · doi:10.48550/arxiv.2307.06844
This paper proposes exploiting the common sense knowledge learned by large language models to perform zero-shot reasoning about crimes given textual descriptions of surveillance videos. We show that when video is (manually) converted to high quality textual descriptions, large language models are capable of detecting and classifying crimes with state-of-the-art performance using only zero-shot reasoning. However, existing automated video-to-text approaches are unable to generate video descriptions of sufficient quality to support reasoning (garbage video descriptions into the large language model, garbage out).