Yuki Mitsufuji
- The Principles of Diffusion Models
2025/10/24 by Chieh-Hsin Lai, Lai, Chieh-Hsin, Yang Song +7 · 21 voices · 11 citations
Computer Science · #cs.LG #cs.AI #cs.GR
- Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
2023/10/01 by Dongjun Kim, Chieh-Hsin Lai, Kim, Dongjun +15 · 137 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Machine Learning and Data Classification
- A Survey on Diffusion Models for Inverse Problems
2024/09/30 by Giannis Daras, Hyungjin Chung, Daras, Giannis +13 · 89 citations
Computer Science · Mathematics · #Advanced Mathematical Modeling in Engineering #Numerical methods in inverse problems #Differential Equations and Numerical Methods
- Manifold Preserving Guided Diffusion
2023/11/28 by Yutong He, He, Yutong, Naoki Murata +19 · 43 citations
Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Model Reduction and Neural Networks
- STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events
2023/06/15 by Kazuki Shimada, Shimada, Kazuki, Archontis Politis +21 · 33 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Animal Vocal Communication and Behavior #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
2022/05/16 by Yuhta Takida, Takashi Shibuya, Takida, Yuhta +17 · 20 citations
Computer Science · #AI in cancer detection #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech and Audio Processing
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
2024/12/19 by Ho Kei Cheng, Cheng, Ho Kei, Masato Ishii +9 · 50 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
2020/10/29 by Kazuki Shimada, Yuichiro Koyama, Shimada, Kazuki +7 · 15 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
2022/06/04 by Archontis Politis, Kazuki Shimada, Politis, Archontis +17 · 13 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Noise Effects and Management #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
2024/05/27 by Junyoung Seo, Kazumi Fukuda, Seo, Junyoung +15 · 20 citations
Computer Science · #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis #Image Processing and 3D Reconstruction
- GibbsDDRM: A Partially Collapsed Gibbs Sampler for Solving Blind Inverse Problems with Denoising Diffusion Restoration
2023/01/30 by Naoki Murata, Koichi Saito, Murata, Naoki +11 · 11 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Jump Your Steps: Optimizing Sampling Schedule of Discrete Diffusion Models
2024/10/10 by Yong-Hyun Park, Chieh-Hsin Lai, Park, Yong-Hyun +7 · 21 citations
Mathematics · #Statistical Methods and Inference
- Automatic Piano Transcription with Hierarchical Frequency-Time Transformer
2023/07/10 by Keisuke Toyama, Toyama, Keisuke, Taketo Akama +9 · 11 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
2022/12/14 by Hao‐Wen Dong, Dong, Hao-Wen, Naoya Takahashi +7 · 10 citations
Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Recursive speech separation for unknown number of speakers
2019/04/05 by Naoya Takahashi, Parthasaarathy Sudarsanam, Takahashi, Naoya +5 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
2024/02/09 by Yixiao Zhang, Yukara Ikemiya, Zhang, Yixiao +13 · 13 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
- Automatic music mixing with deep learning and out-of-domain data
2022/08/24 by Marco A. Martínez-Ramírez, Martínez-Ramírez, Marco A., Wei‐Hsiang Liao +9 · 10 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
- HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
2023/12/31 by Yuhta Takida, Takida, Yuhta, Yukara Ikemiya +19 · 10 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Image and Signal Denoising Methods #AI in cancer detection #Cancer-related molecular mechanisms research
- Distillation of Discrete Diffusion through Dimensional Correlations
2024/10/11 by Satoshi Hayakawa, Yuhta Takida, Hayakawa, Satoshi +7 · 15 citations
Engineering · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Numerical Analysis (math.NA) #Process Optimization and Integration
- Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
2024/07/19 by Roser Batlle-Roca, Batlle-Roca, Roser, Liao, Wei-Hsiang +6 · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects
2022/11/04 by Junghyun Koo, Marco A. Martinez-Ramirez, Koo, Junghyun +9 · 10 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher
2024/05/23 by Dongjun Kim, Chieh-Hsin Lai, Kim, Dongjun +13 · 9 citations
Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Educational Technology and Assessment #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- D3Net: Densely connected multidilated DenseNet for music source separation
2020/10/05 by Naoya Takahashi, Yuki Mitsufuji, Takahashi, Naoya +1 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Hierarchical disentangled representation learning for singing voice conversion
2021/01/18 by Naoya Takahashi, Mayank Singh, Takahashi, Naoya +3 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- FP-Diffusion: Improving Score-based Diffusion Models by Enforcing the Underlying Score Fokker-Planck Equation
2022/10/09 by Chieh-Hsin Lai, Lai, Chieh-Hsin, Yuhta Takida +9 · 5 citations
Computer Science · Mathematics · Physics and Astronomy · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Mathematical Biology Tumor Growth #Model Reduction and Neural Networks
- SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
2024/06/25 by Marco Comunità, Zhi Zhong, Comunità, Marco +17 · 10 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Unsupervised vocal dereverberation with diffusion-based generative models
2022/11/08 by Koichi Saito, Naoki Murata, Saito, Koichi +11 · 4 citations
Computer Science · Engineering · #Speech and Audio Processing #Music and Audio Processing #Acoustic Wave Phenomena Research
- A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
2024/09/26 by Masato Ishii, Ishii, Masato, Akio Hayakawa +5 · 7 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies
- PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives
2023/05/03 by Silin Gao, Gao, Silin, Beatriz Borges +13 · 4 citations
Computer Science · #AI in Service Interactions #Computation and Language (cs.CL) #FOS: Computer and information sciences #Persona Design and Applications #Topic Modeling
- Distortion Audio Effects: Learning How to Recover the Clean Signal
2022/02/03 by Johannes Imort, Giorgio Fabbro, Imort, Johannes +9 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
2024/11/23 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reservoir Engineering and Simulation Methods #cs.AI #cs.CV #cs.LG
- TraSCE: Trajectory Steering for Concept Erasure
2024/12/10 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies
- SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer
2023/01/30 by Yuhta Takida, Takida, Yuhta, Masaaki Imaizumi +10 · 5 citations
Computer Science · #AI in cancer detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #Machine Learning (cs.LG)
- Zero- and Few-shot Sound Event Localization and Detection
2023/09/17 by Kazuki Shimada, Kengo Uchida, Shimada, Kazuki +11 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- The Sound Demixing Challenge 2023 \unicodex2013 Cinematic Demixing Track
2023/08/14 by Stefan Uhlich, Uhlich, Stefan, Giorgio Fabbro +31 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
2024/05/28 by Koichi Saito, Saito, Koichi, D. S. Kim +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
- Preventing Oversmoothing in VAE via Generalized Variance Parameterization
2021/02/17 by Yuhta Takida, Takida, Yuhta, Wei‐Hsiang Liao +9 · 3 citations
Computer Science · #AI in cancer detection #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
- Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
2024/01/01 by Michele Mancusi, Mancusi, Michele, Yurii Halychanskyi +20 · 1 voice · 3 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.AI #cs.IR #cs.SD #eess.AS #electronic engineering #information engineering
- 30+ Years of Source Separation Research: Achievements and Future Challenges
2025/01/21 by Shoko Araki, Nobutaka Ito, Araki, Shoko +9 · 8 citations
Engineering · #Audio and Speech Processing (eess.AS) #Drilling and Well Engineering #FOS: Computer and information sciences #FOS: Electrical engineering #Hydraulic Fracturing and Reservoir Analysis #Microfluidic and Capillary Electrophoresis Applications #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
- Aligning Text-to-Music Evaluation with Human Preferences
2025/03/20 by Yichen Huang, Huang, Yichen, Zachary Novack +13 · 10 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
- Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
2024/03/28 by Yutong He, Alexander Robey, He, Yutong +17 · 4 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mathematics, Computing, and Information Processing #Multimedia Communication and Technology #Video Analysis and Summarization
- Sound Event Localization and Detection Using Activity-Coupled Cartesian DOA Vector and RD3net
2020/06/22 by Kazuki Shimada, Naoya Takahashi, Shimada, Kazuki +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
2023/05/18 by Hao Shi, Shi, Hao, Kazuki Shimada +15 · 3 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- D2USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
2025/04/08 by Jisang Han, Honggyu An, Han, Jisang +16 · 7 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics and Sensor-Based Localization
- Searching For Music Mixing Graphs: A Pruning Approach
2024/06/03 by Sungho Lee, Lee, Sungho, Marco A. Martínez-Ramírez +11 · 6 citations
Arts and Humanities · Computer Science · #Digital Humanities and Scholarship #FOS: Computer and information sciences #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
- Extending Audio Masked Autoencoders Toward Audio Restoration
2023/05/11 by Zhi Zhong, Zhong, Zhi, Hao Shi +13 · 4 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering
- Music Foundation Model as Generic Booster for Music Downstream Tasks
2024/11/02 by WeiHsiang Liao, Liao, WeiHsiang, Yuhta Takida +29 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion
2024/10/18 by Nguyễn Hoàng Bắc, Nguyen, Bac, Lai, Chieh-Hsin +10 · 4 citations
Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Statistical Methods and Inference
- OpenMU: Your Swiss Army Knife for Music Understanding
2024/10/21 by Mengjie Zhao, Zhi Zhong, Zhao, Mengjie +13 · 5 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
2023/09/06 by Takashi Shibuya, Shibuya, Takashi, Yuhta Takida +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
2025/04/15 by Xinlei Niu, Niu, Xinlei, Kin Wai Cheuk +19 · 5 citations
Computer Science · #Artificial Intelligence in Games #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
2025/05/26 by Yigitcan Özer, Özer, Yigitcan, Woosung Choi +9 · 6 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- DiffuCOMET: Contextual Commonsense Knowledge Diffusion
2024/02/26 by Silin Gao, Mete Ismayilzada, Gao, Silin +9 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Neural Networks and Applications #Semantic Web and Ontologies #Topic Modeling
- Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
2024/05/23 by Shiqi Yang, Zhi Zhong, Yang, Shiqi +11 · 3 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Variable Bitrate Residual Vector Quantization for Audio Coding
2024/10/08 by Yunkee Chae, Woosung Choi, Chae, Yunkee +19 · 4 citations
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #Digital Filter Design and Implementation #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- DiffVox: A Differentiable Model for Capturing and Analysing Vocal Effects Distributions
2025/04/20 by Chin-Yun Yu, Marco A. Martínez-Ramírez, Yu, Chin-Yun +12 · 1 voice · 5 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Animal Vocal Communication and Behavior #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.SD #eess.AS #electronic engineering #information engineering
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
2021/06/21 by Kazuki Shimada, Shimada, Kazuki, Naoya Takahashi +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Source Mixing and Separation Robust Audio Steganography
2021/10/11 by Naoya Takahashi, Mayank Kumar Singh, Takahashi, Naoya +3 · 1 citation
Computer Science · #Advanced Steganography and Watermarking Techniques #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- VinaBench: Benchmark for Faithful and Consistent Visual Narratives
2025/03/26 by Silin Gao, Sheryl Mathew, Gao, Silin +16 · 2 voices · 2 citations
Arts and Humanities · Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #FOS: Computer and information sciences #Multimedia Communication and Technology #Video Analysis and Summarization #cs.AI #cs.CL #cs.CV
- MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial Training
2024/06/04 by Kengo Uchida, Uchida, Kengo, Takashi Shibuya +10 · 2 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Human Pose and Action Recognition
- Robust One-Shot Singing Voice Conversion
2022/10/20 by Naoya Takahashi, Takahashi, Naoya, Mayank Singh +3 · 1 citation
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
2025/04/27 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +15 · 5 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
- LLM2Fx-Tools: Tool Calling For Music Post-Production
2025/12/01 by Seungheon Doh, Doh, Seungheon, Junghyun Koo +11 · 3 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
- Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
2025/06/16 by Junyoung Seo, Jisang Han, Seo, Junyoung +21 · 4 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation
- GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
2024/08/06 by Sungho Lee, Lee, Sungho, Marco A. Martínez-Ramírez +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- HumanGif: Single-View Human Diffusion with Generative Prior
2025/02/17 by Shoukang Hu, Takuya Narihira, Hu, Shoukang +9 · 4 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation
- Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric
2024/04/30 by Toshimitsu Uesaka, Uesaka, Toshimitsu, Taiji Suzuki +9 · 2 citations
Arts and Humanities · Psychology · #EFL/ESL Teaching and Learning #FOS: Computer and information sciences #Innovative Teaching and Learning Methods #Machine Learning (cs.LG)
- G2D2: Gradient-Guided Discrete Diffusion for Inverse Problem Solving
2024/10/09 by Naoki Murata, Murata, Naoki, Chieh-Hsin Lai +11 · 2 citations
Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Numerical methods in inverse problems #Photoacoustic and Ultrasonic Imaging
- Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
2025/10/06 by Satoshi Hayakawa, Yuhta Takida, Hayakawa, Satoshi +7 · 2 citations
Computer Science · Medicine · #Generative Adversarial Networks and Image Synthesis #Advanced Neuroimaging Techniques and Applications #Medical Image Segmentation Techniques
- On the Language Encoder of Contrastive Cross-modal Models
2023/10/20 by Mengjie Zhao, Zhao, Mengjie, Junya Ono +17 · 1 citation
Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Subtitles and Audiovisual Media
- Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association
2023/10/02 by Qiyu Wu, Wu, Qiyu, Mengjie Zhao +11 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques
- MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
2024/03/15 by Hao Tan, Kin Wai Cheuk, Tan, Hao Hao +7 · 1 citation
Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Speech Recognition and Synthesis
- Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
2024/06/22 by Yuhua Chen, Chen, Yu-Hua, Woosung Choi +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
2024/06/17 by Hiromi Wakaki, Wakaki, Hiromi, Yuki Mitsufuji +13 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Persona Design and Applications
- CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
2025/01/06 by Yuanhong Chen, Chen, Yuanhong, Kazuki Shimada +9 · 2 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks
2021/10/13 by Boyu Chen, Chen, Bo-Yu, Wei‐Han Hsu +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Bellman Diffusion: Generative Modeling as Learning a Linear Operator in the Distribution Space
2024/10/02 by Yangming Li, Chieh-Hsin Lai, Li, Yangming +7 · 1 citation
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mathematical and Theoretical Analysis #Neural Networks and Applications #Statistical and Computational Modeling
- MeanFlow Transformers with Representation Autoencoders
2025/11/17 by Zheyuan Hu, Chieh-Hsin Lai, Hu, Zheyuan +7 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
- Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection
2021/10/13 by Yuichiro Koyama, Koyama, Yuichiro, Kazuhide Shigemi +13 · 1 citation
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
2024/10/08 by M. Jehanzeb Mirza, Mengjie Zhao, Mirza, M. Jehanzeb +27 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- Cross-Modal Learning for Music-to-Music-Video Description Generation
2025/03/14 by Zhuoyuan Mao, Mao, Zhuoyuan, Mengjie Zhao +11 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
- Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
2025/06/27 by R. Oguz Araz, Araz, R. Oguz, Guillem Cortès +11 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- C3G: Learning Compact 3D Representations with 2K Gaussians
2025/12/03 by Honggyu An, Jaewoo Jung, An, Honggyu +19 · 1 citation
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
2025/10/08 by Jiaben Chen, Chen, Jiaben, Ailing Zeng +18 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing
- SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
2025/05/22 by Zhi Zhong, Akira Takahashi, Zhong, Zhi +9 · 4 citations
Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Generative Adversarial Networks and Image Synthesis
- Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
2024/10/01 by Saurav Jha, Jha, Saurav, Shiqi Yang +17 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Authorship Attribution and Profiling #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing
- Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
2025/12/14 by Weihan Xu, K. y. Cheng, Xu, Weihan +23 · 1 citation
Computer Science · #Video Analysis and Summarization #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing
- Spectral Prior for Reducing Exposure Bias in Diffusion Models
2026/07/24 by Yuya Kobayashi, Masato Ishii, Yuhta Takida +2
Computer Science · #cs.CV