Tagliasacchi, Marco
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Comanici, Gheorghe, Eric Bieber +6844 · 8 voices · 1374 citations
#cs.CL #cs.AI
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024/03/08 by Gemini Robotics Team, Gemini Team, Petko Georgiev +2277 · 4 voices · 559 citations
Computer Science · #Semantic Web and Ontologies
- MuChoMusic dataset
2023/01/26 by Andrea Agostinelli, Timo I. Denk, Agostinelli, Andrea +23 · 1 voice · 67 citations
Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Speech Recognition and Synthesis #cs.LG #cs.SD #eess.AS
- AudioPaLM: A Large Language Model That Can Speak and Listen
2023/06/22 by Paul K. Rubenstein, Chulayuth Asawaroengchai, Rubenstein, Paul K. +57 · 1 voice · 52 citations
#cs.CL #cs.AI #cs.SD #eess.AS #stat.ML
- SoundStream: An End-to-End Neural Audio Codec
2021/07/07 by Zeghidour, Neil, Luebs, Alejandro, Omran, Ahmed +2 · 170 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- AudioLM: a Language Modeling Approach to Audio Generation
2022/09/07 by Zalán Borsos, Raphaël Marinier, Borsos, Zalán +18 · 94 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
2023/02/07 by Eugene Kharitonov, Kharitonov, Eugene, Damien Vincent +15 · 27 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- SEANet: A Multi-modal Speech Enhancement Network
2020/09/04 by Tagliasacchi, Marco, Li, Yunpeng, Misiunas, Karolis +1 · 13 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- SoundStorm: Efficient Parallel Audio Generation
2023/05/16 by Zalán Borsos, Borsos, Zalán, Matt Sharifi +9 · 16 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- LEAF: A Learnable Frontend for Audio Classification
2021/01/21 by Neil Zeghidour, Olivier Teboul, Zeghidour, Neil +5 · 10 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Disentangling speech from surroundings with neural embeddings
2022/03/29 by Omran, Ahmed, Zeghidour, Neil, Borsos, Zalán +3 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- SpeechPainter: Text-conditioned Speech Inpainting
2022/02/15 by Borsos, Zalán, Sharifi, Matt, Tagliasacchi, Marco · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Text-Driven Separation of Arbitrary Sounds
2022/04/12 by Kevin Kilgour, Kilgour, Kevin, Beat Gfeller +9 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Real-time Speech Frequency Bandwidth Extension
2020/10/21 by Li, Yunpeng, Tagliasacchi, Marco, Rybakov, Oleg +2 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Learning to Denoise Historical Music
2020/08/05 by Li, Yunpeng, Gfeller, Beat, Tagliasacchi, Marco +1 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
- MicAugment: One-shot Microphone Style Transfer
2020/10/19 by Borsos, Zalán, Li, Yunpeng, Gfeller, Beat +1 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
2023/03/23 by Jenrungrot, Teerapat, Chinen, Michael, Kleijn, W. Bastiaan +4 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
2023/08/21 by Erdogan, Hakan, Wisdom, Scott, Chang, Xuankai +4 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- MAD Speech: Measures of Acoustic Diversity of Speech
2024/04/16 by Matthieu Futeral, Andrea Agostinelli, Futeral, Matthieu +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering