vix.ing · top · new · best · stats · spec

Petridis, Stavros

  1. End-to-end Audiovisual Speech Recognition
    2018/02/18 by Stavros Petridis, Themos Stafylakis, Petridis, Stavros +9 · 1 voice · 5 citations
    #cs.CV
  2. Lips Don't Lie: A Generalisable and Robust Approach to Face Forgery\n Detection
    2020/12/14 by Alexandros Haliassos, Haliassos, Alexandros, Konstantinos Vougioukas +5 · 41 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
  3. Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation
    2023/01/06 by Stypułkowski, Michał, Vougioukas, Konstantinos, He, Sen +3 · 28 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  4. Leveraging Real Talking Faces via Self-Supervision for Robust Forgery Detection
    2022/01/18 by Alexandros Haliassos, Haliassos, Alexandros, Rodrigo Mira +5 · 16 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis
  5. End-to-end Audio-visual Speech Recognition with Conformers
    2021/02/12 by Ma, Pingchuan, Petridis, Stavros, Pantic, Maja · 14 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  6. EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars
    2024/04/29 by Drobyshev, Nikita, Casademunt, Antoni Bigata, Vougioukas, Konstantinos +3 · 22 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Lipreading using Temporal Convolutional Networks
    2020/01/23 by Martinez, Brais, Ma, Pingchuan, Petridis, Stavros +1 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture
    2018/09/28 by Petridis, Stavros, Stafylakis, Themos, Ma, Pingchuan +2 · 7 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. Jointly Learning Visual and Auditory Speech Representations from Raw Data
    2022/12/12 by Haliassos, Alexandros, Ma, Pingchuan, Mira, Rodrigo +2 · 8 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Sound (cs.SD)
  10. Large Language Models are Strong Audio-Visual Speech Recognition Learners
    2024/09/18 by Cappellazzo, Umberto, Kim, Minsu, Chen, Honglie +5 · 15 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  11. Realistic Speech-Driven Facial Animation with GANs
    2019/06/14 by Vougioukas, Konstantinos, Petridis, Stavros, Pantic, Maja · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  12. End-to-End Speech-Driven Facial Animation with Temporal GANs
    2018/05/23 by Konstantinos Vougioukas, Stavros Petridis, Vougioukas, Konstantinos +3 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Image and Video Processing (eess.IV) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  13. TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch
    2023/10/27 by Hwang, Jeff, Hira, Moto, Chen, Caroline +21 · 6 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  14. BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition
    2024/04/02 by Haliassos, Alexandros, Zinonos, Andreas, Mira, Rodrigo +2 · 6 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  15. End-To-End Visual Speech Recognition With LSTMs
    2017/01/20 by Petridis, Stavros, Li, Zuwei, Pantic, Maja · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  16. Towards Pose-invariant Lip-Reading
    2019/11/14 by Shiyang Cheng, Pingchuan Ma, Cheng, Shiyang +11 · 2 citations
    Computer Science · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research #Speech and Audio Processing
  17. Towards Practical Lipreading with Distilled and Efficient Models
    2020/07/13 by Pingchuan Ma, Ma, Pingchuan, Brais Martínez +5 · 2 citations
    Computer Science · Neuroscience · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hand Gesture Recognition Systems #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Tactile and Sensory Interactions
  18. Domain Adversarial Neural Networks for Dysarthric Speech Recognition
    2020/10/07 by Woszczyk, Dominika, Petridis, Stavros, Millard, David · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  19. LiRA: Learning Visual Speech Representations from Audio through Self-supervision
    2021/06/16 by Ma, Pingchuan, Mira, Rodrigo, Petridis, Stavros +2 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  20. SVTS: Scalable Video-to-Speech Synthesis
    2022/05/04 by Mira, Rodrigo, Haliassos, Alexandros, Petridis, Stavros +2 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  21. LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders
    2022/11/20 by Rodrigo Mira, Mira, Rodrigo, Buye Xu +11 · 2 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Machine Learning (cs.LG) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  22. SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision
    2023/03/30 by Xubo Liu, Egor Lakomkin, Liu, Xubo +21 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Face recognition and analysis #Indoor and Outdoor Localization Technologies
  23. RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
    2024/07/10 by Chen, Honglie, Mira, Rodrigo, Petridis, Stavros +1 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  24. KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
    2025/03/03 by Bigata, Antoni, Stypułkowski, Michał, Mira, Rodrigo +7 · 5 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  25. Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
    2024/11/04 by Alexandros Haliassos, Rodrigo Mira, Haliassos, Alexandros +9 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech and Audio Processing
  26. Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
    2025/03/11 by Minsu Kim, Rodrigo Mira, Kim, Minsu +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  27. Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
    2025/03/09 by Cappellazzo, Umberto, Kim, Minsu, Petridis, Stavros · 6 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  28. Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
    2025/03/08 by Yeo, Jeong Hun, Kim, Minsu, Kim, Chae Won +2 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  29. SS-VAERR: Self-Supervised Apparent Emotional Reaction Recognition from Video
    2022/10/20 by Jegorova, Marija, Petridis, Stavros, Pantic, Maja · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  30. Learning Cross-lingual Visual Speech Representations
    2023/03/14 by Zinonos, Andreas, Haliassos, Alexandros, Ma, Pingchuan +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  31. SparseVSR: Lightweight and Noise Robust Visual Speech Recognition
    2023/07/10 by Fernandez-Lopez, Adriana, Chen, Honglie, Ma, Pingchuan +3 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  32. Hearing Loss Detection from Facial Expressions in One-on-one Conversations
    2024/01/17 by Yufeng Yin, Yin, Yufeng, Ishwarya Ananthabhotla +9 · 1 citation
    Computer Science · Neuroscience · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Hearing Loss and Rehabilitation #Speech and Audio Processing
  33. KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
    2025/05/01 by Bigata, Antoni, Mira, Rodrigo, Bounareli, Stella +4 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  34. Medical records condensation: a roadmap towards healthcare data democratisation
    2023/05/05 by Wang, Yujiang, Thakur, Anshul, Dong, Mingzhi +5 · 1 citation
    #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  35. Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
    2025/05/20 by Umberto Cappellazzo, Cappellazzo, Umberto, Minsu Kim +7 · 3 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  36. FaceCrafter: Identity-Conditional Diffusion with Disentangled Control over Facial Pose, Expression, and Emotion
    2025/05/21 by Mishima, Kazuaki, Casademunt, Antoni Bigata, Petridis, Stavros +2 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences