Takashi Shibuya
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
2022/05/16 by Yuhta Takida, Takashi Shibuya, Takida, Yuhta +17 · 20 citations
Computer Science · #AI in cancer detection #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech and Audio Processing
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
2024/12/19 by Ho Kei Cheng, Cheng, Ho Kei, Masato Ishii +9 · 52 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
2024/05/27 by Junyoung Seo, Seo, Junyoung, Kazumi Fukuda +15 · 20 citations
Computer Science · #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis #Image Processing and 3D Reconstruction
- HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
2023/12/31 by Yuhta Takida, Yukara Ikemiya, Takida, Yuhta +19 · 11 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Image and Signal Denoising Methods #AI in cancer detection #Cancer-related molecular mechanisms research
- SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
2024/06/25 by Marco Comunità, Zhi Zhong, Comunità, Marco +17 · 10 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
2024/09/26 by Masato Ishii, Akio Hayakawa, Ishii, Masato +5 · 7 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies
- Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
2024/11/23 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reservoir Engineering and Simulation Methods #cs.AI #cs.CV #cs.LG
- TraSCE: Trajectory Steering for Concept Erasure
2024/12/10 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies
- Zero- and Few-shot Sound Event Localization and Detection
2023/09/17 by Kazuki Shimada, Shimada, Kazuki, Kengo Uchida +11 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
2024/05/28 by Koichi Saito, D. S. Kim, Saito, Koichi +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
- Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER
2021/10/16 by Dong‐Ho Lee, Akshen Kadakia, Lee, Dong-Ho +17 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
2023/05/18 by Hao Shi, Kazuki Shimada, Shi, Hao +15 · 3 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Extending Audio Masked Autoencoders Toward Audio Restoration
2023/05/11 by Zhi Zhong, Hao Shi, Zhong, Zhi +13 · 4 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering
- BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
2023/09/06 by Takashi Shibuya, Shibuya, Takashi, Yuhta Takida +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Nested Named Entity Recognition via Second-best Sequence Learning and Decoding
2019/09/05 by Takashi Shibuya, Shibuya, Takashi, Eduard Hovy +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
- Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
2024/05/23 by Shiqi Yang, Zhi Zhong, Yang, Shiqi +11 · 3 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
2025/04/27 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +16 · 1 voice · 5 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #cs.CV
- MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial Training
2024/06/04 by Kengo Uchida, Takashi Shibuya, Uchida, Kengo +10 · 2 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Human Pose and Action Recognition
- Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
2025/06/16 by Junyoung Seo, Jisang Han, Seo, Junyoung +21 · 4 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation
- HumanGif: Single-View Human Diffusion with Generative Prior
2025/02/17 by Shoukang Hu, Takuya Narihira, Hu, Shoukang +9 · 4 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation
- On the Language Encoder of Contrastive Cross-modal Models
2023/10/20 by Mengjie Zhao, Junya Ono, Zhao, Mengjie +17 · 1 citation
Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Subtitles and Audiovisual Media
- CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
2025/01/06 by Yuanhong Chen, Kazuki Shimada, Chen, Yuanhong +9 · 2 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
2025/12/14 by Weihan Xu, K. y. Cheng, Xu, Weihan +23 · 1 citation
Computer Science · #Video Analysis and Summarization #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing
- Spectral Prior for Reducing Exposure Bias in Diffusion Models
2026/07/24 by Yuya Kobayashi, Masato Ishii, Yuhta Takida +2
Computer Science · #cs.CV