vix.ing · top · new · best · stats · spec

Yuki Mitsufuji

  1. The Principles of Diffusion Models
    2025/10/24 by Chieh-Hsin Lai, Lai, Chieh-Hsin, Yang Song +7 · 21 voices · 11 citations
    Computer Science · #cs.LG #cs.AI #cs.GR
  2. Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
    2023/10/01 by Dongjun Kim, Chieh-Hsin Lai, Kim, Dongjun +15 · 137 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Machine Learning and Data Classification
  3. A Survey on Diffusion Models for Inverse Problems
    2024/09/30 by Giannis Daras, Hyungjin Chung, Daras, Giannis +13 · 89 citations
    Computer Science · Mathematics · #Advanced Mathematical Modeling in Engineering #Numerical methods in inverse problems #Differential Equations and Numerical Methods
  4. Manifold Preserving Guided Diffusion
    2023/11/28 by Yutong He, He, Yutong, Naoki Murata +19 · 43 citations
    Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Model Reduction and Neural Networks
  5. STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events
    2023/06/15 by Kazuki Shimada, Shimada, Kazuki, Archontis Politis +21 · 33 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Animal Vocal Communication and Behavior #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  6. SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization
    2022/05/16 by Yuhta Takida, Takashi Shibuya, Takida, Yuhta +17 · 20 citations
    Computer Science · #AI in cancer detection #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech and Audio Processing
  7. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
    2024/12/19 by Ho Kei Cheng, Cheng, Ho Kei, Masato Ishii +9 · 50 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
    2020/10/29 by Kazuki Shimada, Yuichiro Koyama, Shimada, Kazuki +7 · 15 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  9. STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
    2022/06/04 by Archontis Politis, Kazuki Shimada, Politis, Archontis +17 · 13 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Noise Effects and Management #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  10. GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
    2024/05/27 by Junyoung Seo, Kazumi Fukuda, Seo, Junyoung +15 · 20 citations
    Computer Science · #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis #Image Processing and 3D Reconstruction
  11. GibbsDDRM: A Partially Collapsed Gibbs Sampler for Solving Blind Inverse Problems with Denoising Diffusion Restoration
    2023/01/30 by Naoki Murata, Koichi Saito, Murata, Naoki +11 · 11 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  12. Jump Your Steps: Optimizing Sampling Schedule of Discrete Diffusion Models
    2024/10/10 by Yong-Hyun Park, Chieh-Hsin Lai, Park, Yong-Hyun +7 · 21 citations
    Mathematics · #Statistical Methods and Inference
  13. Automatic Piano Transcription with Hierarchical Frequency-Time Transformer
    2023/07/10 by Keisuke Toyama, Toyama, Keisuke, Taketo Akama +9 · 11 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  14. CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
    2022/12/14 by Hao‐Wen Dong, Dong, Hao-Wen, Naoya Takahashi +7 · 10 citations
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  15. Recursive speech separation for unknown number of speakers
    2019/04/05 by Naoya Takahashi, Parthasaarathy Sudarsanam, Takahashi, Naoya +5 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  16. MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
    2024/02/09 by Yixiao Zhang, Yukara Ikemiya, Zhang, Yixiao +13 · 13 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
  17. Automatic music mixing with deep learning and out-of-domain data
    2022/08/24 by Marco A. Martínez-Ramírez, Martínez-Ramírez, Marco A., Wei‐Hsiang Liao +9 · 10 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
  18. HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
    2023/12/31 by Yuhta Takida, Takida, Yuhta, Yukara Ikemiya +19 · 10 citations
    Computer Science · Biochemistry, Genetics and Molecular Biology · #Image and Signal Denoising Methods #AI in cancer detection #Cancer-related molecular mechanisms research
  19. Distillation of Discrete Diffusion through Dimensional Correlations
    2024/10/11 by Satoshi Hayakawa, Yuhta Takida, Hayakawa, Satoshi +7 · 15 citations
    Engineering · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Numerical Analysis (math.NA) #Process Optimization and Integration
  20. Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
    2024/07/19 by Roser Batlle-Roca, Batlle-Roca, Roser, Liao, Wei-Hsiang +6 · 12 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  21. Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects
    2022/11/04 by Junghyun Koo, Marco A. Martinez-Ramirez, Koo, Junghyun +9 · 10 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  22. PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher
    2024/05/23 by Dongjun Kim, Chieh-Hsin Lai, Kim, Dongjun +13 · 9 citations
    Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Educational Technology and Assessment #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  23. D3Net: Densely connected multidilated DenseNet for music source separation
    2020/10/05 by Naoya Takahashi, Yuki Mitsufuji, Takahashi, Naoya +1 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  24. Hierarchical disentangled representation learning for singing voice conversion
    2021/01/18 by Naoya Takahashi, Mayank Singh, Takahashi, Naoya +3 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  25. FP-Diffusion: Improving Score-based Diffusion Models by Enforcing the Underlying Score Fokker-Planck Equation
    2022/10/09 by Chieh-Hsin Lai, Lai, Chieh-Hsin, Yuhta Takida +9 · 5 citations
    Computer Science · Mathematics · Physics and Astronomy · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Mathematical Biology Tumor Growth #Model Reduction and Neural Networks
  26. SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
    2024/06/25 by Marco Comunità, Zhi Zhong, Comunità, Marco +17 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  27. Unsupervised vocal dereverberation with diffusion-based generative models
    2022/11/08 by Koichi Saito, Naoki Murata, Saito, Koichi +11 · 4 citations
    Computer Science · Engineering · #Speech and Audio Processing #Music and Audio Processing #Acoustic Wave Phenomena Research
  28. A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
    2024/09/26 by Masato Ishii, Ishii, Masato, Akio Hayakawa +5 · 7 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies
  29. PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives
    2023/05/03 by Silin Gao, Gao, Silin, Beatriz Borges +13 · 4 citations
    Computer Science · #AI in Service Interactions #Computation and Language (cs.CL) #FOS: Computer and information sciences #Persona Design and Applications #Topic Modeling
  30. Distortion Audio Effects: Learning How to Recover the Clean Signal
    2022/02/03 by Johannes Imort, Giorgio Fabbro, Imort, Johannes +9 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  31. Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
    2024/11/23 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reservoir Engineering and Simulation Methods #cs.AI #cs.CV #cs.LG
  32. TraSCE: Trajectory Steering for Concept Erasure
    2024/12/10 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +11 · 1 voice · 6 citations
    Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies
  33. SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer
    2023/01/30 by Yuhta Takida, Takida, Yuhta, Masaaki Imaizumi +10 · 5 citations
    Computer Science · #AI in cancer detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #Machine Learning (cs.LG)
  34. Zero- and Few-shot Sound Event Localization and Detection
    2023/09/17 by Kazuki Shimada, Kengo Uchida, Shimada, Kazuki +11 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  35. The Sound Demixing Challenge 2023 \unicodex2013 Cinematic Demixing Track
    2023/08/14 by Stefan Uhlich, Uhlich, Stefan, Giorgio Fabbro +31 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  36. SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
    2024/05/28 by Koichi Saito, Saito, Koichi, D. S. Kim +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  37. Preventing Oversmoothing in VAE via Generalized Variance Parameterization
    2021/02/17 by Yuhta Takida, Takida, Yuhta, Wei‐Hsiang Liao +9 · 3 citations
    Computer Science · #AI in cancer detection #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
  38. Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
    2024/01/01 by Michele Mancusi, Mancusi, Michele, Yurii Halychanskyi +20 · 1 voice · 3 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.AI #cs.IR #cs.SD #eess.AS #electronic engineering #information engineering
  39. 30+ Years of Source Separation Research: Achievements and Future Challenges
    2025/01/21 by Shoko Araki, Nobutaka Ito, Araki, Shoko +9 · 8 citations
    Engineering · #Audio and Speech Processing (eess.AS) #Drilling and Well Engineering #FOS: Computer and information sciences #FOS: Electrical engineering #Hydraulic Fracturing and Reservoir Analysis #Microfluidic and Capillary Electrophoresis Applications #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  40. Aligning Text-to-Music Evaluation with Human Preferences
    2025/03/20 by Yichen Huang, Huang, Yichen, Zachary Novack +13 · 10 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
  41. Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
    2024/03/28 by Yutong He, Alexander Robey, He, Yutong +17 · 4 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mathematics, Computing, and Information Processing #Multimedia Communication and Technology #Video Analysis and Summarization
  42. Sound Event Localization and Detection Using Activity-Coupled Cartesian DOA Vector and RD3net
    2020/06/22 by Kazuki Shimada, Naoya Takahashi, Shimada, Kazuki +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  43. Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
    2023/05/18 by Hao Shi, Shi, Hao, Kazuki Shimada +15 · 3 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  44. D2USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
    2025/04/08 by Jisang Han, Honggyu An, Han, Jisang +16 · 7 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics and Sensor-Based Localization
  45. Searching For Music Mixing Graphs: A Pruning Approach
    2024/06/03 by Sungho Lee, Lee, Sungho, Marco A. Martínez-Ramírez +11 · 6 citations
    Arts and Humanities · Computer Science · #Digital Humanities and Scholarship #FOS: Computer and information sciences #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
  46. Extending Audio Masked Autoencoders Toward Audio Restoration
    2023/05/11 by Zhi Zhong, Zhong, Zhi, Hao Shi +13 · 4 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering
  47. Music Foundation Model as Generic Booster for Music Downstream Tasks
    2024/11/02 by WeiHsiang Liao, Liao, WeiHsiang, Yuhta Takida +29 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  48. Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion
    2024/10/18 by Nguyễn Hoàng Bắc, Nguyen, Bac, Lai, Chieh-Hsin +10 · 4 citations
    Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Statistical Methods and Inference
  49. OpenMU: Your Swiss Army Knife for Music Understanding
    2024/10/21 by Mengjie Zhao, Zhi Zhong, Zhao, Mengjie +13 · 5 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  50. BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
    2023/09/06 by Takashi Shibuya, Shibuya, Takashi, Yuhta Takida +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  51. SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
    2025/04/15 by Xinlei Niu, Niu, Xinlei, Kin Wai Cheuk +19 · 5 citations
    Computer Science · #Artificial Intelligence in Games #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  52. A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
    2025/05/26 by Yigitcan Özer, Özer, Yigitcan, Woosung Choi +9 · 6 citations
    Computer Science · #Advanced Steganography and Watermarking Techniques #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  53. DiffuCOMET: Contextual Commonsense Knowledge Diffusion
    2024/02/26 by Silin Gao, Mete Ismayilzada, Gao, Silin +9 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Neural Networks and Applications #Semantic Web and Ontologies #Topic Modeling
  54. Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
    2024/05/23 by Shiqi Yang, Zhi Zhong, Yang, Shiqi +11 · 3 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  55. Variable Bitrate Residual Vector Quantization for Audio Coding
    2024/10/08 by Yunkee Chae, Woosung Choi, Chae, Yunkee +19 · 4 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #Digital Filter Design and Implementation #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  56. DiffVox: A Differentiable Model for Capturing and Analysing Vocal Effects Distributions
    2025/04/20 by Chin-Yun Yu, Marco A. Martínez-Ramírez, Yu, Chin-Yun +12 · 1 voice · 5 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Animal Vocal Communication and Behavior #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.SD #eess.AS #electronic engineering #information engineering
  57. Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
    2021/06/21 by Kazuki Shimada, Shimada, Kazuki, Naoya Takahashi +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  58. Source Mixing and Separation Robust Audio Steganography
    2021/10/11 by Naoya Takahashi, Mayank Kumar Singh, Takahashi, Naoya +3 · 1 citation
    Computer Science · #Advanced Steganography and Watermarking Techniques #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  59. VinaBench: Benchmark for Faithful and Consistent Visual Narratives
    2025/03/26 by Silin Gao, Sheryl Mathew, Gao, Silin +16 · 2 voices · 2 citations
    Arts and Humanities · Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #FOS: Computer and information sciences #Multimedia Communication and Technology #Video Analysis and Summarization #cs.AI #cs.CL #cs.CV
  60. MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial Training
    2024/06/04 by Kengo Uchida, Uchida, Kengo, Takashi Shibuya +10 · 2 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Human Pose and Action Recognition
  61. Robust One-Shot Singing Voice Conversion
    2022/10/20 by Naoya Takahashi, Takahashi, Naoya, Mayank Singh +3 · 1 citation
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  62. Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image
    2025/04/27 by Anubhav Jain, Yuya Kobayashi, Jain, Anubhav +15 · 5 citations
    Computer Science · #Advanced Steganography and Watermarking Techniques #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
  63. LLM2Fx-Tools: Tool Calling For Music Post-Production
    2025/12/01 by Seungheon Doh, Doh, Seungheon, Junghyun Koo +11 · 3 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
  64. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
    2025/06/16 by Junyoung Seo, Jisang Han, Seo, Junyoung +21 · 4 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation
  65. GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
    2024/08/06 by Sungho Lee, Lee, Sungho, Marco A. Martínez-Ramírez +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  66. HumanGif: Single-View Human Diffusion with Generative Prior
    2025/02/17 by Shoukang Hu, Takuya Narihira, Hu, Shoukang +9 · 4 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation
  67. Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric
    2024/04/30 by Toshimitsu Uesaka, Uesaka, Toshimitsu, Taiji Suzuki +9 · 2 citations
    Arts and Humanities · Psychology · #EFL/ESL Teaching and Learning #FOS: Computer and information sciences #Innovative Teaching and Learning Methods #Machine Learning (cs.LG)
  68. G2D2: Gradient-Guided Discrete Diffusion for Inverse Problem Solving
    2024/10/09 by Naoki Murata, Murata, Naoki, Chieh-Hsin Lai +11 · 2 citations
    Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Numerical methods in inverse problems #Photoacoustic and Ultrasonic Imaging
  69. Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
    2025/10/06 by Satoshi Hayakawa, Yuhta Takida, Hayakawa, Satoshi +7 · 2 citations
    Computer Science · Medicine · #Generative Adversarial Networks and Image Synthesis #Advanced Neuroimaging Techniques and Applications #Medical Image Segmentation Techniques
  70. On the Language Encoder of Contrastive Cross-modal Models
    2023/10/20 by Mengjie Zhao, Zhao, Mengjie, Junya Ono +17 · 1 citation
    Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Subtitles and Audiovisual Media
  71. Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association
    2023/10/02 by Qiyu Wu, Wu, Qiyu, Mengjie Zhao +11 · 1 citation
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques
  72. MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage
    2024/03/15 by Hao Tan, Kin Wai Cheuk, Tan, Hao Hao +7 · 1 citation
    Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Speech Recognition and Synthesis
  73. Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
    2024/06/22 by Yuhua Chen, Chen, Yu-Hua, Woosung Choi +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  74. ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
    2024/06/17 by Hiromi Wakaki, Wakaki, Hiromi, Yuki Mitsufuji +13 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Persona Design and Applications
  75. CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
    2025/01/06 by Yuanhong Chen, Chen, Yuanhong, Kazuki Shimada +9 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  76. Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks
    2021/10/13 by Boyu Chen, Chen, Bo-Yu, Wei‐Han Hsu +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  77. Bellman Diffusion: Generative Modeling as Learning a Linear Operator in the Distribution Space
    2024/10/02 by Yangming Li, Chieh-Hsin Lai, Li, Yangming +7 · 1 citation
    Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mathematical and Theoretical Analysis #Neural Networks and Applications #Statistical and Computational Modeling
  78. MeanFlow Transformers with Representation Autoencoders
    2025/11/17 by Zheyuan Hu, Chieh-Hsin Lai, Hu, Zheyuan +7 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
  79. Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection
    2021/10/13 by Yuichiro Koyama, Koyama, Yuichiro, Kazuhide Shigemi +13 · 1 citation
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  80. GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
    2024/10/08 by M. Jehanzeb Mirza, Mengjie Zhao, Mirza, M. Jehanzeb +27 · 2 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  81. Cross-Modal Learning for Music-to-Music-Video Description Generation
    2025/03/14 by Zhuoyuan Mao, Mao, Zhuoyuan, Mengjie Zhao +11 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
  82. Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
    2025/06/27 by R. Oguz Araz, Araz, R. Oguz, Guillem Cortès +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  83. C3G: Learning Compact 3D Representations with 2K Gaussians
    2025/12/03 by Honggyu An, Jaewoo Jung, An, Honggyu +19 · 1 citation
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
  84. TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
    2025/10/08 by Jiaben Chen, Chen, Jiaben, Ailing Zeng +18 · 3 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing
  85. SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
    2025/05/22 by Zhi Zhong, Akira Takahashi, Zhong, Zhi +9 · 4 citations
    Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Generative Adversarial Networks and Image Synthesis
  86. Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
    2024/10/01 by Saurav Jha, Jha, Saurav, Shiqi Yang +17 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Authorship Attribution and Profiling #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing
  87. Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
    2025/12/14 by Weihan Xu, K. y. Cheng, Xu, Weihan +23 · 1 citation
    Computer Science · #Video Analysis and Summarization #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing
  88. Spectral Prior for Reducing Exposure Bias in Diffusion Models
    2026/07/24 by Yuya Kobayashi, Masato Ishii, Yuhta Takida +2
    Computer Science · #cs.CV