vix.ing · top · new · best · stats · spec

Wasnik, Pankaj

  1. Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
    2024/12/29 by Ayush Ghadiya, Purbayan Kar, Ghadiya, Ayush +5 · 5 citations
    Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Network Security and Intrusion Detection #Artificial Immune Systems Applications
  2. DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
    2024/06/13 by Neha Sahipjohn, Ashishkumar Gudmalwar, Sahipjohn, Neha +7 · 3 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. M2FNet: Multi-modal Fusion Network for Emotion Recognition in Conversation
    2022/06/05 by Vishal Chudasama, Purbayan Kar, Chudasama, Vishal +9 · 1 citation
    Psychology · Computer Science · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining
  4. VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
    2024/06/12 by Gudmalwar, Ashishkumar, Shah, Nirmesh, Akarsh, Sai +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. Beyond Few-shot Object Detection: A Detailed Survey
    2024/08/26 by Chudasama, Vishal, Sarkar, Hiran, Wasnik, Pankaj +2 · 1 citation
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2.10 #I.4.8 #I.5
  6. EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
    2024/12/29 by Ashishkumar Gudmalwar, Ishan D. Biyani, Gudmalwar, Ashishkumar +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing