vix.ing · top · new · best · stats · spec

Class-attention Video Transformer for Engagement Intensity Prediction

2022/08/12 by Xusheng Ai, Ai, Xusheng, Victor S. Sheng +4 · 3 citations
Computer Science · Neuroscience · #Advanced Technologies in Various Fields #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mind wandering and attention #Online Learning and Analytics

paper · pdf · doi:10.48550/arxiv.2208.07216

openalex publication_date 2022/08/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In order to deal with variant-length long videos, prior works extract multi-modal features and fuse them to predict students' engagement intensity. In this paper, we present a new end-to-end method Class Attention in Video Transformer (CavT), which involves a single vector to process class embedding and to uniformly perform end-to-end learning on variant-length long videos and fixed-length short videos. Furthermore, to address the lack of sufficient samples, we propose a binary-order representatives sampling method (BorS) to add multiple video sequences of each video to augment the training set. BorS+CavT not only achieves the state-of-the-art MSE (0.0495) on the EmotiW-EP dataset, but also obtains the state-of-the-art MSE (0.0377) on the DAiSEE dataset. The code and models have been made publicly available at https://github.com/mountainai/cavt.

Cited by

Related