vix.ing · top · new · best · stats · spec

MorphFader: Enabling Fine-grained Controllable Morphing with Text-to-Audio Models

2024/08/14 by Purnima Kamath, Kamath, Purnima, Chitralekha Gupta +3 · 1 voice · 2 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Human Motion and Animation #Music Technology and Sound Studies #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2408.07260

openalex publication_date 2024/08/14 · arxiv published 2024/08/14 · arxiv updated 2024/08/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced high-quality sounds using text prompts. However, granularly controlling the semantics of the sound, which is necessary for morphing, can be challenging using text. In this paper, we propose MorphFader, a controllable method for morphing sounds generated by disparate prompts using text-to-audio models. By intercepting and interpolating the components of the cross-attention layers within the diffusion process, we can create smooth morphs between sounds generated by different text prompts. Using both objective metrics and perceptual listening tests, we demonstrate the ability of our method to granularly control the semantics in the sound and generate smooth morphs.

Cited by

Discussions

Related