2021/08/04 by Ahmed Ibrahim, Ibrahim, Ahmed, Ayman El-Refai +9
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hand Gesture Recognition Systems #Human-Computer Interaction (cs.HC) #Indoor and Outdoor Localization Technologies #Machine Learning (cs.LG) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2108.02148
openalex publication_date 2021/08/04 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Due to the mass advancement in ubiquitous technologies nowadays, new\npervasive methods have come into the practice to provide new innovative\nfeatures and stimulate the research on new human-computer interactions. This\npaper presents a hand gesture recognition method that utilizes the smartphone's\nbuilt-in speakers and microphones. The proposed system emits an ultrasonic\nsonar-based signal (inaudible sound) from the smartphone's stereo speakers,\nwhich is then received by the smartphone's microphone and processed via a\nConvolutional Neural Network (CNN) for Hand Gesture Recognition. Data\naugmentation techniques are proposed to improve the detection accuracy and\nthree dual-channel input fusion methods are compared. The first method merges\nthe dual-channel audio as a single input spectrogram image. The second method\nadopts early fusion by concatenating the dual-channel spectrograms. The third\nmethod adopts late fusion by having two convectional input branches processing\neach of the dual-channel spectrograms and then the outputs are merged by the\nlast layers. Our experimental results demonstrate a promising detection\naccuracy for the six gestures presented in our publicly available dataset with\nan accuracy of 93.58 % as a baseline.\n