vix.ing · top · new · best · stats · spec

Mohtashami, Amirkeivan

  1. Landmark Attention: Random-Access Infinite Context Length for Transformers
    2023/05/25 by Amirkeivan Mohtashami, Martin Jaggi, Mohtashami, Amirkeivan +1 · 2 voices · 16 citations
    Computer Science · #Topic Modeling #Advanced Neural Network Applications #Machine Learning and Data Classification
  2. DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
    2024/02/04 by Matteo Pagliardini, Amirkeivan Mohtashami, Pagliardini, Matteo +6 · 1 voice · 13 citations
    Computer Science · #Neural Networks and Applications
  3. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
    2024/03/30 by Saleh Ashkboos, Ashkboos, Saleh, Amirkeivan Mohtashami +14 · 105 citations
    Computer Science · Engineering · #Advanced Wireless Communication Techniques #Error Correcting Code Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Optical Network Technologies
  4. MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
    2023/11/27 by Zeming Chen, A. Cano, Chen, Zeming +37 · 44 citations
    Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Topic Modeling
  5. CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
    2023/10/16 by Mohtashami, Amirkeivan, Pagliardini, Matteo, Jaggi, Martin · 11 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Social Learning: Towards Collaborative Learning with Large Language Models
    2023/12/18 by Amirkeivan Mohtashami, Mohtashami, Amirkeivan, Florian Hartmann +11 · 1 voice · 2 citations
    Computer Science · #Privacy-Preserving Technologies in Data #Topic Modeling #cs.CL #cs.LG
  7. Masked Training of Neural Networks with Partial Gradients
    2021/06/16 by Amirkeivan Mohtashami, Mohtashami, Amirkeivan, Martin Jaggi +3 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Stochastic Gradient Optimization Techniques
  8. Special Properties of Gradient Descent with Large Learning Rates
    2022/05/30 by Amirkeivan Mohtashami, Mohtashami, Amirkeivan, Martin Jaggi +3 · 2 citations
    Computer Science · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning and Algorithms #Machine Learning and ELM #Optimization and Control (math.OC) #Stochastic Gradient Optimization Techniques
  9. Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
    2021/03/03 by Stich, Sebastian U., Mohtashami, Amirkeivan, Jaggi, Martin · 1 citation
    #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Parallel #and Cluster Computing (cs.DC)
  10. Characterizing & Finding Good Data Orderings for Fast Convergence of Sequential Gradient Methods
    2022/02/03 by Mohtashami, Amirkeivan, Stich, Sebastian, Jaggi, Martin · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)