vix.ing · top · new · best · stats · spec

Guangzhi Sun

  1. Large language models surpass human experts in predicting neuroscience results
    2024/03/04 by Xiaoliang Luo, Akilles Rechardt, Guangzhi Sun +41 · 5 voices · 36 citations
    Computer Science · Materials Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Machine Learning in Materials Science
  2. SALMONN: Towards Generic Hearing Abilities for Large Language Models
    2023/10/20 by Changli Tang, Tang, Changli, Wenyi Yu +15 · 110 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
    2024/06/22 by Guangzhi Sun, Wenyi Yu, Sun, Guangzhi +17 · 25 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Speech and Audio Processing
  4. Connecting Speech Encoder and Large Language Model for ASR
    2023/09/25 by Wenyi Yu, Yu, Wenyi, Changli Tang +15 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
    2024/09/25 by Siyin Wang, Wenyi Yu, Wang, Siyin +20 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  6. Building Better AI Agents: A Provocation on the Utilisation of Persona in LLM-based Conversational Agents
    2024/05/26 by Guangzhi Sun, Sun, Guangzhi, Xiao Zhan +3 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Persona Design and Applications
  7. CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
    2025/01/24 by Guangzhi Sun, Sun, Guangzhi, 晓慧 战 +7 · 8 citations
    Computer Science · Social Sciences · #Access Control and Trust #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling
  8. Can Large Language Models Understand Spatial Audio?
    2024/06/12 by Changli Tang, Tang, Changli, Wenyi Yu +19 · 6 citations
    Computer Science · Social Sciences · #Audio and Speech Processing (eess.AS) #Computational and Text Analysis Methods #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  9. Can Contextual Biasing Remain Effective with Whisper and GPT-2?
    2023/06/02 by Guangzhi Sun, Sun, Guangzhi, Xianrui Zheng +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  10. Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models
    2023/10/09 by Guangzhi Sun, Wenyi Yu, Sun, Guangzhi +14 · 3 citations
    Computer Science · #Music and Audio Processing #Multimodal Machine Learning Applications #Speech and Audio Processing
  11. Transformer Language Models with LSTM-based Cross-utterance Information Representation
    2021/02/12 by Guangzhi Sun, C. Zhang, Sun, G. +3 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  12. video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
    2025/02/17 by Guangzhi Sun, Sun, Guangzhi, Yudong Yang +11 · 6 citations
    Computer Science · #Multimodal Machine Learning Applications #Music and Audio Processing #Natural Language Processing Techniques
  13. M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
    2024/03/21 by Zhe Chen, Heyang Liu, Chen, Zhe +15 · 2 citations
    Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Subtitles and Audiovisual Media #Video Analysis and Summarization
  14. Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
    2022/05/18 by Guangzhi Sun, Sun, Guangzhi, Chao Zhang +3 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  15. Bayesian WeakS-to-Strong from Text Classification to Generation
    2024/05/24 by Ziyun Cui, Ziyang Zhang, Cui, Ziyun +7 · 2 citations
    Computer Science · #Text and Document Classification Technologies
  16. Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
    2025/11/13 by Yudong Yang, Xuezhen Zhang, Yang, Yudong +14 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #User Authentication and Security Systems #Music and Audio Processing
  17. Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
    2024/10/09 by Changli Tang, Yixuan Li, Tang, Changli +11 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimodal Machine Learning Applications #Video Analysis and Summarization #electronic engineering #information engineering
  18. Matching domain experts by training from scratch on domain knowledge
    2024/05/15 by Xiaoliang Luo, Guangzhi Sun, Luo, Xiaoliang +3 · 3 voices
    #q-bio.NC #cs.AI #cs.CL
  19. SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
    2026/07/19 by Xiaoyu Yang, Xuenan Xu, Wenyi Yu +10
    #eess.AS