vix.ing · top · new · best · stats · spec

Bai, Xinyi

  1. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1393 citations
    #cs.CL #cs.AI
  2. Faithful Persona-based Conversational Dataset Generation with Large Language Models
    2023/12/15 by Pegah Jandaghi, Xianghai Sheng, Jandaghi, Pegah +7 · 11 citations
    Computer Science · #AI in Service Interactions #Computation and Language (cs.CL) #FOS: Computer and information sciences #Innovative Human-Technology Interaction #Machine Learning (cs.LG) #Persona Design and Applications
  3. Towards Rationality in Language and Multimodal Agents: A Survey
    2024/06/01 by Jiang, Bowen, Xie, Yangxinyu, Wang, Xiaomeng +6 · 9 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multiagent Systems (cs.MA)
  4. Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
    2024/01/31 by Jiatong Shi, Shi, Jiatong, Yueqian Lin +14 · 6 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
    2024/09/11 by Wu, Yuning, Shi, Jiatong, Yu, Yifeng +7 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
    2024/09/26 by Epstein, Elliot L., Kaisheng Yao, Yao, Kaisheng +6 · 1 citation
    Computer Science · #Speech and dialogue systems #Natural Language Processing Techniques
  7. SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
    2025/10/03 by C.X. Hao, Hao, Chunbo, Ruibin Yuan +10 · 2 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing