Deepak Narayanan
- On the Opportunities and Risks of Foundation Models
2021/08/16 by Rishi Bommasani, Bommasani, Rishi, Drew A. Hudson +233 · 11 voices · 551 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Topic Modeling
- An Empirical Study of Mamba-based Language Models
2024/06/12 by Roger Waleffe, Waleffe, Roger, Wonmin Byeon +30 · 1 voice · 46 citations
Arts and Humanities · #Language, Linguistics, Cultural Analysis #cs.CL #cs.LG
- Efficient Large-Scale Language Model Training on GPU Clusters Using\n Megatron-LM
2021/04/09 by Deepak Narayanan, Narayanan, Deepak, Mohammad Shoeybi +21 · 111 citations
Computer Science · #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
- Llama-Nemotron: Efficient Reasoning Models
2025/05/02 by Akhiad Bercovich, Itay Levy, Bercovich, Akhiad +224 · 1 voice · 29 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
- PipeDream: Fast and Efficient Pipeline Parallel DNN Training
2018/06/08 by Aaron Harlap, Deepak Narayanan, Harlap, Aaron +11 · 31 citations
Computer Science · #Advanced Neural Network Applications #Distributed #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Parallel #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
- MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
2022/11/29 by Trevor Gale, Deepak Narayanan, Gale, Trevor +5 · 28 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Stochastic Gradient Optimization Techniques #Machine Learning and ELM
- Pretraining Large Language Models with NVFP4
2025/09/29 by NVIDIA, Felix Abecassis, Abecassis, Felix +175 · 4 voices · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads
2020/08/20 by Deepak Narayanan, Keshav Santhanam, Narayanan, Deepak +7 · 18 citations
Computer Science · #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #Parallel #Parallel Computing and Optimization Techniques #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
- MLPerf Training Benchmark
2019/10/02 by Peter Mattson, Mattson, Peter, Christine Cheng +71 · 20 citations
Computer Science · #Advanced Neural Network Applications #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Performance (cs.PF)
- Nemotron-4 15B Technical Report
2024/02/26 by Jupinder Parmar, Parmar, Jupinder, Shrimai Prabhumoye +51 · 1 voice · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
2025/08/20 by NVIDIA, :, Aarti Basant +305 · 24 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Big Data and Digital Economy #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel Computing and Optimization Techniques
- Weld: Rethinking the Interface Between Data-Intensive Applications
2017/09/14 by Shoumik Palkar, James Thomas, Palkar, Shoumik +17 · 1 voice
Computer Science · #Databases (cs.DB) #Distributed #FOS: Computer and information sciences #Parallel #Performance (cs.PF) #and Cluster Computing (cs.DC) #cs.DB #cs.DC #cs.PF
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
2025/12/23 by NVIDIA, Aaron Blakeman, : +623 · 1 voice · 2 citations
#cs.CL #cs.AI #cs.LG