Marco Tagliasacchi
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024/03/08 by Gemini Robotics Team, Gemini Team, Petko Georgiev +2277 · 4 voices · 1048 citations
Computer Science · #Semantic Web and Ontologies
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1438 citations
Computer Science · #cs.CL #cs.AI
- MuChoMusic dataset
2023/01/26 by Andrea Agostinelli, Agostinelli, Andrea, Timo I. Denk +23 · 1 voice · 110 citations
Computer Science · Engineering · #Music Technology and Sound Studies #Music and Audio Processing #Speech Recognition and Synthesis #cs.LG #cs.SD #eess.AS
- AudioPaLM: A Large Language Model That Can Speak and Listen
2023/06/22 by Paul K. Rubenstein, Chulayuth Asawaroengchai, Rubenstein, Paul K. +59 · 1 voice · 90 citations
Computer Science · Engineering · Mathematics · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.AI #cs.CL #cs.SD #eess.AS #stat.ML
- AudioLM: a Language Modeling Approach to Audio Generation
2022/09/07 by Zalán Borsos, Borsos, Zalán, Raphaël Marinier +18 · 166 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
2023/02/07 by Eugene Kharitonov, Damien Vincent, Kharitonov, Eugene +15 · 44 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- SoundStorm: Efficient Parallel Audio Generation
2023/05/16 by Zalán Borsos, Matt Sharifi, Borsos, Zalán +9 · 33 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- LEAF: A Learnable Frontend for Audio Classification
2021/01/21 by Neil Zeghidour, Zeghidour, Neil, Olivier Teboul +5 · 18 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Real-time Speech Frequency Bandwidth Extension
2020/10/21 by Yunpeng Li, Li, Yunpeng, Marco Tagliasacchi +7 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Text-Driven Separation of Arbitrary Sounds
2022/04/12 by Kevin Kilgour, Beat Gfeller, Kilgour, Kevin +9 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Disentangling speech from surroundings with neural embeddings
2022/03/29 by Ahmed Omran, Omran, Ahmed, Neil Zeghidour +9 · 6 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
2023/08/21 by Hakan Erdoğan, Erdogan, Hakan, Scott Wisdom +11 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
2023/03/23 by Teerapat Jenrungrot, Michael Chinen, Jenrungrot, Teerapat +11 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Self-supervised audio representation learning for mobile devices
2019/05/24 by Marco Tagliasacchi, Tagliasacchi, Marco, Beat Gfeller +5 · 1 citation
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Learning to Denoise Historical Music
2020/08/05 by Yunpeng Li, Beat Gfeller, Li, Yunpeng +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
- CycleGAN-Based Unpaired Speech Dereverberation
2022/03/29 by Hannah Muckenhirn, Aleksandr Safin, Muckenhirn, Hannah +11 · 1 citation
Computer Science · Psychology · #Speech and Audio Processing #Speech Recognition and Synthesis #Phonetics and Phonology Research
- MAD Speech: Measures of Acoustic Diversity of Speech
2024/04/16 by Matthieu Futeral, Futeral, Matthieu, Andrea Agostinelli +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering