vix.ing · top · new · best · stats · spec

Marco Tagliasacchi

  1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
    2024/03/08 by Gemini Robotics Team, Gemini Team, Petko Georgiev +2277 · 4 voices · 1048 citations
    Computer Science · #Semantic Web and Ontologies
  2. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1438 citations
    Computer Science · #cs.CL #cs.AI
  3. MuChoMusic dataset
    2023/01/26 by Andrea Agostinelli, Agostinelli, Andrea, Timo I. Denk +23 · 1 voice · 110 citations
    Computer Science · Engineering · #Music Technology and Sound Studies #Music and Audio Processing #Speech Recognition and Synthesis #cs.LG #cs.SD #eess.AS
  4. AudioPaLM: A Large Language Model That Can Speak and Listen
    2023/06/22 by Paul K. Rubenstein, Chulayuth Asawaroengchai, Rubenstein, Paul K. +59 · 1 voice · 90 citations
    Computer Science · Engineering · Mathematics · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.AI #cs.CL #cs.SD #eess.AS #stat.ML
  5. AudioLM: a Language Modeling Approach to Audio Generation
    2022/09/07 by Zalán Borsos, Borsos, Zalán, Raphaël Marinier +18 · 166 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  6. Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
    2023/02/07 by Eugene Kharitonov, Damien Vincent, Kharitonov, Eugene +15 · 44 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  7. SoundStorm: Efficient Parallel Audio Generation
    2023/05/16 by Zalán Borsos, Matt Sharifi, Borsos, Zalán +9 · 33 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. LEAF: A Learnable Frontend for Audio Classification
    2021/01/21 by Neil Zeghidour, Zeghidour, Neil, Olivier Teboul +5 · 18 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  9. Real-time Speech Frequency Bandwidth Extension
    2020/10/21 by Yunpeng Li, Li, Yunpeng, Marco Tagliasacchi +7 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Text-Driven Separation of Arbitrary Sounds
    2022/04/12 by Kevin Kilgour, Beat Gfeller, Kilgour, Kevin +9 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. Disentangling speech from surroundings with neural embeddings
    2022/03/29 by Ahmed Omran, Omran, Ahmed, Neil Zeghidour +9 · 6 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  12. TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
    2023/08/21 by Hakan Erdoğan, Erdogan, Hakan, Scott Wisdom +11 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
    2023/03/23 by Teerapat Jenrungrot, Michael Chinen, Jenrungrot, Teerapat +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  14. Self-supervised audio representation learning for mobile devices
    2019/05/24 by Marco Tagliasacchi, Tagliasacchi, Marco, Beat Gfeller +5 · 1 citation
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  15. Learning to Denoise Historical Music
    2020/08/05 by Yunpeng Li, Beat Gfeller, Li, Yunpeng +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
  16. CycleGAN-Based Unpaired Speech Dereverberation
    2022/03/29 by Hannah Muckenhirn, Aleksandr Safin, Muckenhirn, Hannah +11 · 1 citation
    Computer Science · Psychology · #Speech and Audio Processing #Speech Recognition and Synthesis #Phonetics and Phonology Research
  17. MAD Speech: Measures of Acoustic Diversity of Speech
    2024/04/16 by Matthieu Futeral, Futeral, Matthieu, Andrea Agostinelli +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering