vix.ing · top · new · best · stats · spec

Bodapati, Sravan

  1. SpeechVerse: A Large-scale Generalizable Audio Language Model
    2024/05/14 by Nilaksh Das, Das, Nilaksh, Saket Dingliwal +29 · 19 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale
    2022/12/18 by Bansal, Hritik, Gopalakrishnan, Karthik, Dingliwal, Saket +3 · 9 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  3. Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
    2024/10/03 by Elangovan, Aparna, Xu, Lei, Ko, Jongwoo +4 · 13 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC)
  4. The Amazon Nova Family of Models: Technical Report and Model Card
    2025/03/17 by Amazon AGI, AGI, Amazon, Aayush Shah +867 · 22 citations
    Computer Science · Engineering · #3D Modeling in Geospatial Applications #Artificial Intelligence (cs.AI) #BIM and Construction Integration #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model-Driven Software Engineering Techniques
  5. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
    2025/10/06 by Yufeng Du, Minyang Tian, Du, Yufeng +16 · 17 citations
    Computer Science · #Semantic Web and Ontologies #Algorithms and Data Compression #Natural Language Processing Techniques
  6. Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR
    2023/04/18 by Xilai Li, Li, Xilai, Goeric Huybrechts +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
    2024/10/26 by Sullam Jeoung, Jeoung, Sullam, Goeric Huybrechts +7 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics
  8. Multilingual Contextual Adapters To Improve Custom Word Recognition In Low-resource Languages
    2023/07/03 by Devang Kulshreshtha, Saket Dingliwal, Kulshreshtha, Devang +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  9. Retrieve and Copy: Scaling ASR Personalization to Large Catalogs
    2023/11/14 by Sai Muralidhar Jayanthi, Devang Kulshreshtha, Jayanthi, Sai Muralidhar +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  10. Multimodal Semi-supervised Learning Framework for Punctuation Prediction\n in Conversational Speech
    2020/08/03 by Monica Sunkara, Sunkara, Monica, Srikanth Ronanki +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Natural Language Processing Techniques
  11. Best of Both Worlds: Robust Accented Speech Recognition with Adversarial\n Transfer Learning
    2021/03/09 by Nilaksh Das, Das, Nilaksh, Sravan Bodapati +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Mask The Bias: Improving Domain-Adaptive Generalization of CTC-based ASR with Internal Language Model Estimation
    2023/05/05 by Nilaksh Das, Das, Nilaksh, Monica Sunkara +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  13. Generalized zero-shot audio-to-intent classification
    2023/11/04 by Veera Raghavendra Elluru, Devang Kulshreshtha, Elluru, Veera Raghavendra +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  14. Zero-Shot Reinforcement Learning with Deep Attention Convolutional Neural Networks
    2020/01/02 by Şahika Genç, Sunil Mallya, Genc, Sahika +7 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
  15. SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
    2024/10/12 by Jongwoo Ko, Saket Dingliwal, Ko, Jongwoo +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling