vix.ing · top · new · best · stats · spec

Dong Yu

  1. R-Zero: Self-Evolving Reasoning LLM from Zero Data
    2025/08/07 by Chengsong Huang, Wenhao Yu, Huang, Chengsong +15 · 12 voices · 50 citations
    #cs.LG #cs.AI #cs.CL
  2. Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
    2012/10/19 by Geoffrey Hinton, Geoffrey E. Hinton, Li Deng +12 · 400 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  3. Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
    2024/12/30 by Xingyu Chen, Jiahao Xu, Chen, Xingyu +27 · 2 voices · 207 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL #semigroups and automata theory
  4. Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
    2025/01/30 by Yue Wang, Qiuzhi Liu, Wang, Yue +27 · 4 voices · 77 citations
    Computer Science · Social Sciences · #Artificial Intelligence in Law #Digital Rights Management and Security #cs.CL
  5. Scaling Synthetic Data Creation with 1,000,000,000 Personas
    2024/06/28 by Tao Ge, Xin Chan, Ge, Tao +11 · 3 voices · 130 citations
    Computer Science · #Innovative Human-Technology Interaction #Persona Design and Applications #cs.CL #cs.LG
  6. Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
    2024/04/18 by Ye Tian, Baolin Peng, Tian, Ye +12 · 2 voices · 28 citations
    Computer Science · Social Sciences · #Artificial Intelligence in Law #Open Education and E-Learning #Semantic Web and Ontologies #cs.CL #cs.LG
  7. Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
    2025/03/31 by Yi Su, Su Yi, Dian Yu +15 · 3 voices · 63 citations
    Computer Science · Materials Science · #Machine Learning in Materials Science #Multimodal Machine Learning Applications #Topic Modeling #cs.CL
  8. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
    2024/01/25 by Hongliang He, Wenlin Yao, He, Hongliang +13 · 108 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  9. Achieving Human Parity in Conversational Speech Recognition
    2016/10/17 by Wayne Xiong, Xiong, W., Jasha Droppo +13 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Diffsound: Discrete Diffusion Model for Text-to-sound Generation
    2022/07/20 by Dongchao Yang, Jianwei Yu, Yang, Dongchao +11 · 38 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  11. MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
    2023/11/15 by Fuxiao Liu, Liu, Fuxiao, Xiaoyang Wang +13 · 45 citations
    Computer Science · #Speech and dialogue systems
  12. DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
    2025/04/15 by Zhiwei He, He, Zhiwei, Tian Liang +27 · 91 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning and Data Classification #Multimodal Machine Learning Applications #Topic Modeling
  13. An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and\n Separation
    2020/08/21 by Daniel Michelsanti, Michelsanti, Daniel, Zheng‐Hua Tan +11 · 19 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Blind Source Separation Techniques
  14. InFoBench: Evaluating Instruction Following Ability in Large Language Models
    2024/01/07 by Yiwei Qin, Kaiqiang Song, Qin, Yiwei +17 · 32 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  15. A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation
    2023/07/08 by Neeraj Varshney, Varshney, Neeraj, Wenlin Yao +7 · 27 citations
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Text Readability and Simplification #Topic Modeling
  16. Time Domain Audio Visual Speech Separation
    2019/04/07 by Jian Wu, Yong Xu, Wu, Jian +11 · 15 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  17. MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization
    2024/02/18 by Zhiyu Yang, Yang, Zhiyu, Zihan Zhou +23 · 31 citations
    Decision Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Scientific Computing and Data Management
  18. A Fast and Accurate One-Stage Approach to Visual Grounding
    2019/08/18 by Zhengyuan Yang, Yang, Zhengyuan, Boqing Gong +9 · 15 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  19. Multi-talker Speech Separation with Utterance-level Permutation\n Invariant Training of Deep Recurrent Neural Networks
    2017/03/18 by Morten Kolbæk, Dong Yu, Kolbæk, Morten +5 · 13 citations
    Computer Science · Psychology · #Speech and Audio Processing #Speech Recognition and Synthesis #Phonetics and Phonology Research
  20. Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
    2023/11/15 by Wenhao Yu, Hongming Zhang, Yu, Wenhao +9 · 30 citations
    Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  21. Neural Target Speech Extraction: An overview
    2023/05/01 by Katerina Zmolikova, Kateřina Žmolíková, Marc Delcroix +5 · 22 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  22. RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
    2024/10/03 by Siru Ouyang, Ouyang, Siru, Wenhao Yu +15 · 30 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Scientific Computing and Data Management #Software Engineering (cs.SE) #Software Engineering Research #Software Testing and Debugging Techniques
  23. One Token to Fool LLM-as-a-Judge
    2025/07/11 by Yulai Zhao, Haolin Liu, Zhao, Yulai +11 · 2 voices · 21 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CL #cs.LG
  24. LASER: LLM Agent with State-Space Exploration for Web Navigation
    2023/09/15 by Kaixin Ma, Ma, Kaixin, Hongming Zhang +8 · 17 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  25. Dense X Retrieval: What Retrieval Granularity Should We Use?
    2023/12/11 by Tong Chen, Hongwei Wang, Chen, Tong +13 · 18 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  26. FD-GAN: Generative Adversarial Networks with Fusion-discriminator for Single Image Dehazing
    2020/01/20 by Dong Yu, Yihao Liu, Dong, Yu +7 · 9 citations
    Computer Science · Engineering · #Advanced Image Fusion Techniques #Advanced Image Processing Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Enhancement Techniques
  27. SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts
    2021/05/07 by Zhao You, You, Zhao, Shulin Feng +5 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  28. Deep Learning based Multi-Source Localization with Source Splitting and its Effectiveness in Multi-Talker Speech Recognition
    2021/02/16 by Aswin Shanmugam Subramanian, Chao Weng, Subramanian, Aswin Shanmugam +7 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  29. High-Fidelity 3D Digital Human Head Creation from RGB-D Selfies
    2020/10/12 by Linchao Bao, Bao, Linchao, Xiangkai Lin +21 · 6 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Graphics (cs.GR)
  30. EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
    2024/09/17 by Jiarui Hai, Yong Xu, Hai, Jiarui +11 · 17 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  31. Preference Alignment Improves Language Model-Based TTS
    2024/09/19 by Jinchuan Tian, Chunlei Zhang, Tian, Jinchuan +11 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies #Speech and dialogue systems
  32. Recurrent Chunking Mechanisms for Long-Text Machine Reading Comprehension
    2020/05/16 by Hongyu Gong, Gong, Hongyu, Yelong Shen +7 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  33. Deep Extractor Network for Target Speaker Recovery From Single Channel Speech Mixtures
    2018/07/24 by Jun Wang, Wang, Jun, Jie Chen +11 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  34. DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
    2022/01/28 by Songxiang Liu, Dan Su, Liu, Songxiang +3 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  35. Prompt-guided Precise Audio Editing with Diffusion Models
    2024/05/11 by Manjie Xu, Chenxing Li, Xu, Manjie +9 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  36. Robust Disentangled Variational Speech Representation Learning for Zero-shot Voice Conversion
    2022/03/30 by Jiachen Lian, Chunlei Zhang, Lian, Jiachen +3 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  37. Dialogue-Based Relation Extraction
    2020/04/17 by Dian Yu, Yu, Dian, Kai Sun +5 · 4 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  38. Video-to-Audio Generation with Hidden Alignment
    2024/07/10 by Manjie Xu, Xu, Manjie, Chenxing Li +12 · 9 citations
    Computer Science · #Music Technology and Sound Studies
  39. Replay and Synthetic Speech Detection with Res2net Architecture
    2020/10/28 by Xu Li, Li, Xu, Na Li +11 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  40. OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
    2024/10/25 by Hongliang He, Wenlin Yao, He, Hongliang +13 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques #Semantic Web and Ontologies
  41. Sandglasset: A Light Multi-Granularity Self-attentive Network For Time-Domain Speech Separation
    2021/03/01 by Max W. Y. Lam, Lam, Max W. Y., Jun Wang +5 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  42. Make-A-Voice: Unified Voice Synthesis With Discrete Representation
    2023/05/30 by Rongjie Huang, Huang, Rongjie, Chunlei Zhang +17 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  43. FAST-RIR: Fast neural diffuse room impulse response generator
    2021/10/07 by Anton Ratnarajah, Shixiong Zhang, Ratnarajah, Anton +9 · 4 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Indoor and Outdoor Localization Technologies
  44. The Trickle-down Impact of Reward (In-)consistency on RLHF
    2023/09/28 by Lingfeng Shen, Sihao Chen, Shen, Lingfeng +13 · 6 citations
    Computer Science · #Reinforcement Learning in Robotics
  45. WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
    2025/04/23 by Tianqing Fang, Hongming Zhang, Fang, Tianqing +11 · 16 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Domain Adaptation and Few-Shot Learning
  46. MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
    2024/05/29 by Zhenwen Liang, Liang, Zhenwen, Dian Yu +11 · 7 citations
    Social Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Mathematics Education and Teaching Techniques
  47. Tuberous Sclerosis Complex Protein 2-Independent Activation of mTORC1 by Human Cytomegalovirus pUL38
    2015/05/14 by Yadan Bai, Baoqin Xuan, Haiyan Liu +4 · 23 citations
    Biochemistry, Genetics and Molecular Biology · Medicine · #PI3K/AKT/mTOR signaling in cancer #Tuberous Sclerosis Complex Research #Polyomavirus and related diseases
  48. Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning
    2020/03/09 by Rongzhi Gu, Gu, Rongzhi, Shixiong Zhang +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  49. ADL-MVDR: All deep learning MVDR beamformer for target speech separation
    2020/08/16 by Zhuohuang Zhang, Yong Xu, Zhang, Zhuohuang +9 · 4 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Speech Recognition and Synthesis
  50. Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
    2020/07/03 by Liwei Wang, Wang, Liwei, Jing Huang +9 · 3 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  51. Comprehensive Image Captioning via Scene Graph Decomposition
    2020/07/23 by Yiwu Zhong, Zhong, Yiwu, Liwei Wang +7 · 3 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  52. Multi-channel Multi-frame ADL-MVDR for Target Speech Separation
    2020/12/24 by Zhuohuang Zhang, Zhang, Zhuohuang, Yong Xu +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  53. Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
    2024/10/09 by Xiyao Wang, Wang, Xiyao, Song, Linfeng +12 · 8 citations
    Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  54. Importance-based Neuron Allocation for Multilingual Neural Machine Translation
    2021/07/14 by Wanying Xie, Yang Feng, Xie, Wanying +5 · 3 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
  55. Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
    2020/06/11 by Xu Li, Na Li, Li, Xu +13 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Wireless Signal Modulation Classification #electronic engineering #information engineering
  56. Audio-visual Recognition of Overlapped speech for the LRS2 dataset
    2020/01/06 by Jianwei Yu, Shi-Xiong Zhang, Yu, Jianwei +17 · 3 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  57. Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
    2024/12/21 by Zhisong Zhang, Yan Wang, Zhang, Zhisong +13 · 8 citations
    Computer Science · #Topic Modeling #Machine Learning in Healthcare #Explainable Artificial Intelligence (XAI)
  58. LAE: Language-Aware Encoder for Monolingual and Multilingual ASR
    2022/06/05 by Jinchuan Tian, Tian, Jinchuan, Jianwei Yu +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  59. Automatic Prosody Annotation with Pre-Trained Text-Speech Model
    2022/06/16 by Ziqian Dai, Jianwei Yu, Dai, Ziqian +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  60. Toward Unifying Text Segmentation and Long Document Summarization
    2022/10/28 by Sangwoo Cho, Cho, Sangwoo, Kaiqiang Song +7 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  61. Conceptual and Unbiased Reasoning in Language Models
    2024/03/30 by Ben Zhou, Hongming Zhang, Zhou, Ben +13 · 5 citations
    Computer Science · #Natural Language Processing Techniques
  62. Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
    2024/06/17 by Zhihan Zhang, Zhang, Zhihan, Ge, Tao +12 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning
  63. Improving Pre-Trained Multilingual Models with Vocabulary Expansion
    2019/09/26 by Hai Wang, Dian Yu, Wang, Hai +7 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  64. NeuralKalman: A Learnable Kalman Filter for Acoustic Echo Cancellation
    2023/01/29 by Yixuan Zhang, Zhang, Yixuan, Meng Yu +7 · 3 citations
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  65. MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
    2025/07/08 by Yucheng Shi, Shi, Yucheng, Wenhao Yu +13 · 13 citations
    Computer Science · Psychology · #Reinforcement Learning in Robotics #Social Robot Interaction and HRI #Multimodal Machine Learning Applications
  66. Neural Spatio-Temporal Beamformer for Target Speech Separation
    2020/05/08 by Yong Xu, Xu, Yong, Meng Yu +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  67. MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning
    2020/05/11 by Jie Lei, Lei, Jie, Liwei Wang +9 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Video Analysis and Summarization
  68. Advancing Multi-talker ASR Performance with Large Language Models
    2024/08/30 by Mohan Shi, Shi, Mohan, Zengrui Jin +14 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  69. LiteSearch: Efficacious Tree Search for LLM
    2024/06/29 by Ante Wang, Linfeng Song, Wang, Ante +13 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies
  70. Token-level Adaptive Training for Neural Machine Translation
    2020/10/09 by Shuhao Gu, Jinchao Zhang, Gu, Shuhao +11 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  71. WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation
    2020/11/18 by Zhaoheng Ni, Ni, Zhaoheng, Yong Xu +11 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
  72. MinT: Boosting Generalization in Mathematical Reasoning via Multi-View Fine-Tuning
    2023/07/16 by Zhenwen Liang, Dian Yu, Liang, Zhenwen +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  73. Highway Long Short-Term Memory RNNs for Distant Speech Recognition
    2015/10/30 by Yu Zhang, Zhang, Yu, Guoguo Chen +9 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  74. MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment
    2021/04/02 by Meng Yu, Chunlei Zhang, Yu, Meng +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  75. Skills-in-Context Prompting: Unlocking Compositionality in Large Language Models
    2023/08/01 by Jiaao Chen, Xiaoman Pan, Chen, Jiaao +11 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  76. Abstraction-of-Thought Makes Language Models Better Reasoners
    2024/06/18 by Ruixin Hong, Hong, Ruixin, Hongming Zhang +7 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies
  77. Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
    2020/10/30 by Aswin Shanmugam Subramanian, Subramanian, Aswin Shanmugam, Chao Weng +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  78. Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
    2025/05/20 by Xingyu Chen, Wang, Mengru, Chen, Xingyu +25 · 9 citations
    Computer Science · #Cognitive Science and Mapping #AI-based Problem Solving and Planning
  79. Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
    2021/11/29 by Brian Yan, Chunlei Zhang, Yan, Brian +15 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  80. Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching
    2018/12/22 by Chih‐Kuan Yeh, Yeh, Chih-Kuan, Jianshu Chen +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  81. LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
    2025/02/19 by Hao Zhang, Zhang, Hao, Weiwei Li +9 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multi-Agent Systems and Negotiation #Service-Oriented Architecture and Web Services #Speech and dialogue systems #electronic engineering #information engineering
  82. Efficient Zero-shot Event Extraction with Context-Definition Alignment
    2022/11/09 by Hongming Zhang, Zhang, Hongming, Wenlin Yao +3 · 2 citations
    Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #Expert finding and Q&A systems #FOS: Computer and information sciences #Topic Modeling
  83. Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
    2025/08/01 by Tianqing Fang, Fang, Tianqing, Zhisong Zhang +28 · 16 citations
    Computer Science · Medicine · #Topic Modeling #Multimodal Machine Learning Applications #Artificial Intelligence in Healthcare and Education
  84. Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
    2025/09/09 by Tong Zheng, Zheng, Tong, Hongming Zhang +16 · 19 citations
    Computer Science · #Evolutionary Algorithms and Applications
  85. OASum: Large-Scale Open Domain Aspect-based Summarization
    2022/12/19 by Xianjun Yang, Kaiqiang Song, Yang, Xianjun +11 · 2 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #Wikis in Education and Collaboration
  86. STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
    2024/09/13 by Yong Ren, Ren, Yong, Chenxing Li +11 · 4 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies
  87. Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
    2024/10/02 by Mengzhao Jia, Jia, Mengzhao, Wenhao Yu +15 · 5 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Multimodal Machine Learning Applications
  88. Cognitive Kernel: An Open-source Agent System towards Generalist Autopilots
    2024/09/16 by Hongming Zhang, Xiaoman Pan, Zhang, Hongming +9 · 4 citations
    Computer Science · #Neural Networks and Applications
  89. Faithful Question Answering with Monte-Carlo Planning
    2023/05/04 by Ruixin Hong, Hongming Zhang, Hong, Ruixin +7 · 2 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  90. VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
    2025/05/28 by C. Zhang, Kaixin Ma, Zhang, Ce +15 · 7 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications
  91. Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training
    2017/07/19 by Yanmin Qian, Qian, Yanmin, Xuankai Chang +3 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  92. uSee: Unified Speech Enhancement and Editing with Conditional Diffusion Models
    2023/10/02 by Muqiao Yang, Chunlei Zhang, Yang, Muqiao +11 · 2 citations
    Computer Science · Health Professions · #Speech and Audio Processing #Speech Recognition and Synthesis #Infant Health and Development
  93. TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation
    2021/03/31 by Helin Wang, Wang, Helin, Bo Wu +17 · 2 citations
    Computer Science · Neuroscience · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Music and Audio Processing
  94. Deep Audio Zooming: Beamwidth-Controllable Neural Beamformer
    2023/11/22 by Meng Yu, Dong Yu, Yu, Meng +1 · 2 citations
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
  95. The Human Cytomegalovirus Protein pUL38 Suppresses Endoplasmic Reticulum Stress-Mediated Cell Death Independently of Its Ability To Induce mTORC1 Activation
    2011/06/29 by Zhikang Qian, Baoqin Xuan, Nathaniel Gualberto +3 · 26 citations
    Medicine · Biochemistry, Genetics and Molecular Biology · #Cytomegalovirus and herpesvirus research #Endoplasmic Reticulum Stress and Disease #RNA regulation and disease
  96. Router-Tuning: A Simple and Effective Approach for Enabling Dynamic-Depth in Transformers
    2024/10/17 by Shwai He, He, Shwai, Tao Ge +9 · 3 citations
    Engineering · #Modular Robots and Swarm Intelligence
  97. Mixup-breakdown: a consistency training method for improving generalization of speech separation models
    2019/10/28 by Max W. Y. Lam, Jun Wang, Lam, Max W. Y. +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  98. Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams
    2020/06/20 by Huirong Huang, Zhiyong Wu, Huang, Huirong +21 · 1 citation
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Face recognition and analysis
  99. SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
    2024/02/15 by Yebowen Hu, Kaiqiang Song, Hu, Yebowen +11 · 2 citations
    Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Sports Analytics and Performance
  100. Audio-visual Multi-channel Integration and Recognition of Overlapped Speech
    2020/11/16 by Jianwei Yu, Yu, Jianwei, Shi-Xiong Zhang +15 · 1 citation
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
  101. A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression
    2024/12/23 by Chenlong Deng, Deng, Chenlong, Zhisong Zhang +11 · 4 citations
    Computer Science · #Parallel Computing and Optimization Techniques #Embedded Systems Design Techniques
  102. Semantic Role Labeling Guided Multi-turn Dialogue ReWriter
    2020/10/03 by Kun Xu, Xu, Kun, Haochen Tan +11 · 1 citation
    Computer Science · #Topic Modeling #Speech and dialogue systems #Natural Language Processing Techniques
  103. WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
    2025/04/16 by Zhisong Zhang, Zhang, Zhisong, Tianqing Fang +11 · 6 citations
    Computer Science · #Multi-Agent Systems and Negotiation #Peer-to-Peer Network Technologies #Spam and Phishing Detection
  104. Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks
    2021/01/13 by Max W. Y. Lam, Lam, Max W. Y., Jun Wang +5 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  105. Towards Robust Speaker Verification with Target Speaker Enhancement
    2021/03/16 by Chunlei Zhang, Meng Yu, Zhang, Chunlei +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  106. MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation
    2021/04/17 by Xiyun Li, Li, Xiyun, Yong Xu +11 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  107. Structural Information Preserving for Graph-to-Text Generation
    2021/02/12 by Linfeng Song, Ante Wang, Song, Linfeng +11 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  108. Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition
    2021/06/08 by Max W. Y. Lam, Jun Wang, Lam, Max W. Y. +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  109. Bilateral Denoising Diffusion Models
    2021/08/26 by Max W. Y. Lam, Lam, Max W. Y., Jun Wang +7 · 1 citation
    Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  110. Towards Diverse and Efficient Audio Captioning via Diffusion Models
    2024/09/14 by Manjie Xu, Chenxing Li, Xu, Manjie +11 · 3 citations
    Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech and Audio Processing #Subtitles and Audiovisual Media
  111. Towards Improved Zero-shot Voice Conversion with Conditional DSVAE
    2022/05/11 by Jiachen Lian, Chunlei Zhang, Lian, Jiachen +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  112. FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
    2026/06/08 by Yan Wang, Qifan Zhang, Jiachen Yu +12 · 2 voices · 1 citation
    #cs.LG #cs.AI
  113. Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models
    2022/10/28 by Xiaoman Pan, Wenlin Yao, Pan, Xiaoman +9 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  114. Discover, Explanation, Improvement: An Automatic Slice Detection Framework for Natural Language Processing
    2022/11/08 by Wenyue Hua, Lifeng Jin, Hua, Wenyue +9 · 1 citation
    Computer Science · #Topic Modeling #Software Engineering Research #Natural Language Processing Techniques
  115. DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
    2025/05/29 by Ziyin Zhang, Zhang, Ziyin, Jiahao Xu +23 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Logic, programming, and type systems #Multi-Agent Systems and Negotiation #Software Engineering Research
  116. Towards Unified All-Neural Beamforming for Time and Frequency Domain Speech Separation
    2022/12/16 by Rongzhi Gu, Shixiong Zhang, Gu, Rongzhi +5 · 1 citation
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  117. PIVOINE: Instruction Tuning for Open-world Information Extraction
    2023/05/24 by Keming Lu, Xiaoman Pan, Lu, Keming +9 · 1 citation
    Computer Science · Decision Sciences · #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  118. Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
    2025/09/18 by Yujun Zhou, Zhou, Yujun, Haolin Liu +16 · 11 citations
    Computer Science · Medicine · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Multimodal Machine Learning Applications
  119. Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional Operations
    2023/05/24 by James Y. Huang, Wenlin Yao, Huang, James Y. +9 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Natural Language Processing Techniques #Topic Modeling
  120. Fine-Grained Self-Endorsement Improves Factuality and Reasoning
    2024/02/23 by Ante Wang, Linfeng Song, Wang, Ante +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Software Engineering Research #Topic Modeling
  121. Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine Translation
    2023/10/20 by Wenyu Guo, Guo, Wenyu, Qingkai Fang +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  122. Modeling Fluency and Faithfulness for Diverse Neural Machine Translation
    2019/11/30 by Yang Feng, Feng, Yang, Wanying Xie +11 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
  123. When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
    2024/06/17 by Yebowen Hu, Hu, Yebowen, Kaiqiang Song +13 · 1 citation
    Economics, Econometrics and Finance · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Digital Games and Media #FOS: Computer and information sciences #Sports Analytics and Performance #Sports, Gender, and Society
  124. Self-Teaching Machines to Read and Comprehend with Large-Scale Multi-Subject Question-Answering Data
    2021/02/01 by Dian Yu, Yu, Dian, Kai Sun +5 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  125. TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs
    2023/11/09 by Shuyi Xie, Xie, Shuyi, Wenlin Yao +25 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  126. Video-to-Audio Generation with Fine-grained Temporal Semantics
    2024/09/23 by Y. Hu, Hu, Yuchen, Yu Gu +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Video Analysis and Summarization #electronic engineering #information engineering
  127. Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
    2025/02/16 by Ante Wang, Linfeng Song, Wang, Ante +15 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  128. SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
    2025/09/22 by Wei Tan, Shun Lei, Tan, Wei +15 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  129. UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
    2025/09/19 by Chenlong Deng, Deng, Chenlong, Zhisong Zhang +15 · 1 citation
    Computer Science · #Advanced Data Compression Techniques #Algorithms and Data Compression #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis
  130. Field-Induced Dissociation Reveals Excitonic Long-Range Photocarrier Transport in Bulk-Insulating Bi2Se3 Nanoribbons
    2026/07/27 by Rodrigo Becerra Silva, Xiang Yi, Ziyi Song +1
    #cond-mat.mes-hall #cond-mat.str-el