2017/10/02 by Marco Godi, Paolo Rota, Godi, Marco +3
Computer Science · Economics, Econometrics and Finance · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Sports Analytics and Performance #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.1710.00568
openalex publication_date 2017/10/02 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28
Highlights in a sport video are usually referred as actions that stimulate\nexcitement or attract attention of the audience. A big effort is spent in\ndesigning techniques which find automatically highlights, in order to\nautomatize the otherwise manual editing process. Most of the state-of-the-art\napproaches try to solve the problem by training a classifier using the\ninformation extracted on the tv-like framing of players playing on the game\npitch, learning to detect game actions which are labeled by human observers\naccording to their perception of highlight. Obviously, this is a long and\nexpensive work. In this paper, we reverse the paradigm: instead of looking at\nthe gameplay, inferring what could be exciting for the audience, we directly\nanalyze the audience behavior, which we assume is triggered by events happening\nduring the game. We apply deep 3D Convolutional Neural Network (3D-CNN) to\nextract visual features from cropped video recordings of the supporters that\nare attending the event. Outputs of the crops belonging to the same frame are\nthen accumulated to produce a value indicating the Highlight Likelihood (HL)\nwhich is then used to discriminate between positive (i.e. when a highlight\noccurs) and negative samples (i.e. standard play or time-outs). Experimental\nresults on a public dataset of ice-hockey matches demonstrate the effectiveness\nof our method and promote further research in this new exciting direction.\n