vix.ing · top · new · best · stats · spec

Dey, Nolan

  1. Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
    2023/04/06 by Nolan Dey, Gurpreet Gosal, Dey, Nolan +13 · 1 voice · 17 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Machine Learning and Data Classification
  2. Don't be lazy: CompleteP enables compute-efficient deep transformers
    2025/05/02 by Dey, Nolan, Bin Zhang, Zhang, Bin Claire +14 · 1 voice · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Parallel Computing and Optimization Techniques
  3. BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
    2023/09/20 by Nolan Dey, Daria Soboleva, Dey, Nolan +25 · 2 voices · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.AI #cs.CL #cs.LG
  4. Sparse maximal update parameterization: A holistic approach to sparse training dynamics
    2024/05/24 by Nolan Dey, Shane Bergsma, Dey, Nolan +3 · 1 voice · 2 citations
    Computer Science · Physics and Astronomy · #Advanced Vision and Imaging #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Speech and Audio Processing #cs.LG
  5. Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
    2025/02/21 by Bergsma, Shane, Dey, Nolan, Gosal, Gurpreet +3 · 14 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
  6. Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
    2025/05/19 by Bergsma, Shane, Dey, Nolan, Gosal, Gurpreet +3 · 15 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  7. Scaling with Collapse: Efficient and Predictable Training of LLM Families
    2025/09/29 by Bergsma, Shane, Zhang, Bin Claire, Dey, Nolan +3 · 4 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)