vix.ing · top · new · best · stats

Recurrent CNN for 3D Gaze Estimation using Appearance and Shape Cues

2018/05/08 by Cristina Palmero, Palmero, Cristina, Javier Selva +6 · 5 citations
Computer Science · Social Sciences · #Advanced Computing and Algorithms #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gaze Tracking and Assistive Technology #Hand Gesture Recognition Systems #cs.CV

paper · pdf · doi:10.48550/arxiv.1805.03064

Proc. of British Machine Vision Conference (BMVC), BMVC 2018. Errata: in pg.5 the camera matrices of the transformation matrix W should be interchanged (correct version: W=C_n*M*(C_o)^-1)

openalex publication_date 2018/05/08 · arxiv created 2018/09/17 · arxiv updated 2018/09/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Gaze behavior is an important non-verbal cue in social signal processing and human-computer interaction. In this paper, we tackle the problem of person- and head pose-independent 3D gaze estimation from remote cameras, using a multi-modal recurrent convolutional neural network (CNN). We propose to combine face, eyes region, and face landmarks as individual streams in a CNN to estimate gaze in still images. Then, we exploit the dynamic nature of gaze by feeding the learned features of all the frames in a sequence to a many-to-one recurrent module that predicts the 3D gaze vector of the last frame. Our multi-modal static solution is evaluated on a wide range of head poses and gaze directions, achieving a significant improvement of 14.6% over the state of the art on EYEDIAP dataset, further improved by 4% when the temporal modality is included.

Citations

Cited by

Related