vix.ing · top · new · best · stats · spec

Se Jin Park

  1. SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory
    2022/11/02 by Se Jin Park, Minsu Kim, Park, Se Jin +7 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Image and Video Processing (eess.IV) #Speech and Audio Processing #electronic engineering #information engineering
  2. Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
    2024/06/12 by Se Jin Park, Park, Se Jin, Chae Won Kim +11 · 9 citations
    Computer Science · #Speech and dialogue systems #Multi-Agent Systems and Negotiation
  3. MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
    2025/03/14 by Jeong Hun Yeo, Hyeongseop Rha, Yeo, Jeong Hun +5 · 11 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  4. Long-Form Speech Generation with Spoken Language Models
    2024/12/24 by Se Jin Park, Park, Se Jin, Julian Salazar +9 · 7 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Natural Language Processing Techniques
  5. Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model
    2023/10/23 by Joanna Hong, Hong, Joanna, Se Jin Park +3 · 2 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Subtitles and Audiovisual Media #electronic engineering #information engineering
  6. Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
    2022/04/04 by Minsu Kim, Kim, Minsu, Joanna Hong +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  7. Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
    2023/05/31 by Se Jin Park, Minsu Kim, Park, Se Jin +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Image and Video Processing (eess.IV) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  8. Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
    2024/01/18 by Minsu Kim, Kim, Minsu, Jeong Hun Yeo +6 · 1 citation
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Subtitles and Audiovisual Media #Video Surveillance and Tracking Methods #electronic engineering #information engineering
  9. AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
    2024/12/23 by Se Jin Park, Yeonju Kim, Park, Se Jin +7 · 1 citation
    Arts and Humanities · #Media Influence and Health