Mohtashami, Amirkeivan
- Landmark Attention: Random-Access Infinite Context Length for Transformers
2023/05/25 by Amirkeivan Mohtashami, Martin Jaggi, Mohtashami, Amirkeivan +1 · 2 voices · 16 citations
Computer Science · #Topic Modeling #Advanced Neural Network Applications #Machine Learning and Data Classification
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
2024/02/04 by Matteo Pagliardini, Amirkeivan Mohtashami, Pagliardini, Matteo +6 · 1 voice · 13 citations
Computer Science · #Neural Networks and Applications
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
2024/03/30 by Saleh Ashkboos, Ashkboos, Saleh, Amirkeivan Mohtashami +14 · 105 citations
Computer Science · Engineering · #Advanced Wireless Communication Techniques #Error Correcting Code Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Optical Network Technologies
- MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
2023/11/27 by Zeming Chen, A. Cano, Chen, Zeming +37 · 44 citations
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Topic Modeling
- CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
2023/10/16 by Mohtashami, Amirkeivan, Pagliardini, Matteo, Jaggi, Martin · 11 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Social Learning: Towards Collaborative Learning with Large Language Models
2023/12/18 by Amirkeivan Mohtashami, Mohtashami, Amirkeivan, Florian Hartmann +11 · 1 voice · 2 citations
Computer Science · #Privacy-Preserving Technologies in Data #Topic Modeling #cs.CL #cs.LG
- Masked Training of Neural Networks with Partial Gradients
2021/06/16 by Amirkeivan Mohtashami, Mohtashami, Amirkeivan, Martin Jaggi +3 · 2 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Stochastic Gradient Optimization Techniques
- Special Properties of Gradient Descent with Large Learning Rates
2022/05/30 by Amirkeivan Mohtashami, Mohtashami, Amirkeivan, Martin Jaggi +3 · 2 citations
Computer Science · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning and Algorithms #Machine Learning and ELM #Optimization and Control (math.OC) #Stochastic Gradient Optimization Techniques
- Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
2021/03/03 by Stich, Sebastian U., Mohtashami, Amirkeivan, Jaggi, Martin · 1 citation
#Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Parallel #and Cluster Computing (cs.DC)
- Characterizing & Finding Good Data Orderings for Fast Convergence of Sequential Gradient Methods
2022/02/03 by Mohtashami, Amirkeivan, Stich, Sebastian, Jaggi, Martin · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)