vix.ing · top · new · best · stats · spec

Takashi Shibuya

  1. SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
    2022/05/16 by Yuhta Takida, Takashi Shibuya, Takida, Yuhta +17 · 20 citations
    Computer Science · #AI in cancer detection #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech and Audio Processing
  2. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
    2024/12/19 by Ho Kei Cheng, Cheng, Ho Kei, Masato Ishii +9 · 52 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
    2024/05/27 by Junyoung Seo, Seo, Junyoung, Kazumi Fukuda +15 · 20 citations
    Computer Science · #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis #Image Processing and 3D Reconstruction
  4. HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
    2023/12/31 by Yuhta Takida, Yukara Ikemiya, Takida, Yuhta +19 · 11 citations
    Computer Science · Biochemistry, Genetics and Molecular Biology · #Image and Signal Denoising Methods #AI in cancer detection #Cancer-related molecular mechanisms research
  5. SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
    2024/06/25 by Marco Comunità, Zhi Zhong, Comunità, Marco +17 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  6. A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
    2024/09/26 by Masato Ishii, Akio Hayakawa, Ishii, Masato +5 · 7 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies
  7. Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
    2024/11/23 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reservoir Engineering and Simulation Methods #cs.AI #cs.CV #cs.LG
  8. TraSCE: Trajectory Steering for Concept Erasure
    2024/12/10 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
    Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies
  9. Zero- and Few-shot Sound Event Localization and Detection
    2023/09/17 by Kazuki Shimada, Shimada, Kazuki, Kengo Uchida +11 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
    2024/05/28 by Koichi Saito, D. S. Kim, Saito, Koichi +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  11. Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER
    2021/10/16 by Dong‐Ho Lee, Akshen Kadakia, Lee, Dong-Ho +17 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  12. Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
    2023/05/18 by Hao Shi, Kazuki Shimada, Shi, Hao +15 · 3 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. Extending Audio Masked Autoencoders Toward Audio Restoration
    2023/05/11 by Zhi Zhong, Hao Shi, Zhong, Zhi +13 · 4 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering
  14. BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
    2023/09/06 by Takashi Shibuya, Shibuya, Takashi, Yuhta Takida +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. Nested Named Entity Recognition via Second-best Sequence Learning and Decoding
    2019/09/05 by Takashi Shibuya, Shibuya, Takashi, Eduard Hovy +1 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
  16. Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
    2024/05/23 by Shiqi Yang, Zhi Zhong, Yang, Shiqi +11 · 3 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  17. Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
    2025/04/27 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +16 · 1 voice · 5 citations
    Computer Science · #Advanced Steganography and Watermarking Techniques #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #cs.CV
  18. MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial Training
    2024/06/04 by Kengo Uchida, Takashi Shibuya, Uchida, Kengo +10 · 2 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Human Pose and Action Recognition
  19. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
    2025/06/16 by Junyoung Seo, Jisang Han, Seo, Junyoung +21 · 4 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation
  20. HumanGif: Single-View Human Diffusion with Generative Prior
    2025/02/17 by Shoukang Hu, Takuya Narihira, Hu, Shoukang +9 · 4 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation
  21. On the Language Encoder of Contrastive Cross-modal Models
    2023/10/20 by Mengjie Zhao, Junya Ono, Zhao, Mengjie +17 · 1 citation
    Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Subtitles and Audiovisual Media
  22. CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
    2025/01/06 by Yuanhong Chen, Kazuki Shimada, Chen, Yuanhong +9 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  23. Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
    2025/12/14 by Weihan Xu, K. y. Cheng, Xu, Weihan +23 · 1 citation
    Computer Science · #Video Analysis and Summarization #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing
  24. Spectral Prior for Reducing Exposure Bias in Diffusion Models
    2026/07/24 by Yuya Kobayashi, Masato Ishii, Yuhta Takida +2
    Computer Science · #cs.CV