vix.ing · top · new · best · stats · spec

Beidi Chen

  1. Efficient Streaming Language Models with Attention Sinks
    2023/09/29 by Guangxuan Xiao, Yuandong Tian, Xiao, Guangxuan +7 · 5 voices · 417 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
  2. SLIDE : In Defense of Smart Algorithms over Hardware Acceleration for Large-Scale Deep Learning Systems
    2019/03/07 by Beidi Chen, Tharun Medini, Chen, Beidi +9 · 4 voices · 8 citations
    Computer Science · #Advanced Neural Network Applications #Parallel Computing and Optimization Techniques #Machine Learning and Data Classification
  3. GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
    2024/03/06 by Jiawei Zhao, Zhenyu Zhang, Zhao, Jiawei +9 · 5 voices · 70 citations
    Computer Science · #Neural Networks and Applications #Machine Learning and ELM #Advanced Neural Network Applications
  4. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
    2023/03/13 by Ying Sheng, Sheng, Ying, Lianmin Zheng +25 · 2 voices · 96 citations
    Computer Science · Engineering · Materials Science · #Ferroelectric and Negative Capacitance Devices #Machine Learning in Materials Science #Topic Modeling #cs.AI #cs.LG #cs.PF
  5. Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
    2024/04/12 by Xuezhe Ma, Xiaomeng Yang, Ma, Xuezhe +17 · 3 voices · 4 citations
    #cs.LG #cs.CL
  6. KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization
    2023/01/01 by Zirui Liu, Jiayi Yuan, Hongye Jin +6 · 1 voice · 51 citations
    Computer Science · #Error Correcting Code Techniques #Interconnection Networks and Systems #Quantum-Dot Cellular Automata #cs.CL #cs.LG #cs.PF
  7. Decentralized Training of Foundation Models in Heterogeneous Environments
    2022/06/02 by Binhang Yuan, Yongjun He, Yuan, Binhang +16 · 1 voice · 12 citations
    Computer Science · #Advanced Neural Network Applications #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Parallel Computing and Optimization Techniques #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC) #cs.DC #cs.LG
  8. LLM Inference Unveiled: Survey and Roofline Model Insights
    2024/02/26 by Zhihang Yuan, Yuzhang Shang, Yuan, Zhihang +24 · 28 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Mathematics, Computing, and Information Processing
  9. Memory Mosaics
    2024/05/10 by Jianyu Zhang, Niklas Nolte, Zhang, Jianyu +7 · 3 voices · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #cs.AI #cs.LG #cs.NE
  10. ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
    2024/10/28 by Sun, Hanshi, Li‐Wen Chang, Chang, Li-Wen +14 · 28 citations
    Computer Science · #Anomaly Detection Techniques and Applications #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Packet Processing and Optimization #Time Series Analysis and Forecasting
  11. Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
    2024/03/05 by Zhenyu Zhang, Runjin Chen, Zhang, Zhenyu +13 · 19 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech and dialogue systems
  12. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
    2024/08/20 by Ranajoy Sadhukhan, Jian Chen, Sadhukhan, Ranajoy +14 · 22 citations
    Computer Science · #Video Analysis and Summarization #Image Retrieval and Classification Techniques #Multimodal Machine Learning Applications
  13. MagicPIG: LSH Sampling for Efficient LLM Generation
    2024/10/21 by Zhuoming Chen, Ranajoy Sadhukhan, Chen, Zhuoming +19 · 21 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Video Analysis and Summarization
  14. SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
    2024/06/04 by Ruslan Svirschevski, Avner May, Svirschevski, Ruslan +9 · 16 citations
    Computer Science · #Computation and Language (cs.CL) #Digital Rights Management and Security #FOS: Computer and information sciences #Mathematics, Computing, and Information Processing
  15. Monarch: Expressive Structured Matrices for Efficient and Accurate Training
    2022/04/01 by Tri Dao, Beidi Chen, Dao, Tri +17 · 9 citations
    Computer Science · #Advanced Neural Network Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Parallel Computing and Optimization Techniques
  16. TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
    2024/04/18 by Hanshi Sun, Zhuoming Chen, Sun, Hanshi +7 · 14 citations
    Computer Science · #Algorithms and Data Compression #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel Computing and Optimization Techniques
  17. Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer
    2023/05/25 by Yuandong Tian, Tian, Yuandong, Yiping Wang +5 · 10 citations
    Computer Science · Materials Science · Physics and Astronomy · #Neural Networks and Applications #Machine Learning in Materials Science #Model Reduction and Neural Networks
  18. Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
    2021/11/30 by Tri Dao, Beidi Chen, Dao, Tri +11 · 7 citations
    Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  19. HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
    2023/11/20 by Youhe Jiang, Ran Yan, Jiang, Youhe +8 · 11 citations
    Computer Science · Decision Sciences · Engineering · #Distributed #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Modular Robots and Swarm Intelligence #Parallel #Simulation Techniques and Applications #and Cluster Computing (cs.DC)
  20. Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
    2024/02/14 by Harry Dong, Xinyu Yang, Dong, Harry +9 · 10 citations
    Computer Science · #Algorithms and Data Compression #Advanced Data Storage Technologies
  21. Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
    2024/02/19 by Zhuoming Chen, Avner May, Chen, Zhuoming +11 · 10 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Neural Networks and Applications
  22. Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
    2025/06/02 by Haizhong Zheng, Yang Zhou, Zheng, Haizhong +11 · 29 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation
  23. GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
    2025/02/07 by Yang Zhou, Zhou, Yang, Hongyi Liu +7 · 18 citations
    Computer Science · #Digital Rights Management and Security #Multi-Agent Systems and Negotiation
  24. Fine-tuning Language Models over Slow Networks using Activation Compression with Guarantees
    2022/06/02 by Jue Wang, Binhang Yuan, Wang, Jue +13 · 5 citations
    Computer Science · #Stochastic Gradient Optimization Techniques #Machine Learning and ELM #Advanced Neural Network Applications
  25. Compress, Then Prompt: Improving Accuracy-Efficiency Trade-off of LLM Inference with Transferable Prompt
    2023/05/17 by Zhaozhuo Xu, Zirui Liu, Xu, Zhaozhuo +13 · 4 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
  26. Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions
    2023/10/28 by Stefano Massaroli, Michael Poli, Massaroli, Stefano +23 · 5 citations
    Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Signal Processing (eess.SP) #Topic Modeling #electronic engineering #information engineering
  27. APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding
    2025/02/08 by Xinyu Yang, Yang, Xinyu, Tianqi Chen +3 · 8 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Multimodal Machine Learning Applications
  28. On the Surprising Effectiveness of Attention Transfer for Vision Transformers
    2024/11/14 by Alexander C. Li, Yuandong Tian, Li, Alexander C. +7 · 3 voices · 2 citations
    #cs.LG #cs.AI #cs.CV #cs.NE
  29. Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
    2024/05/29 by Minghan Li, Xilun Chen, Ari Holtzman +4 · 1 voice · 1 citation
    Computer Science · #cs.CL
  30. Zeroth-Order Fine-Tuning of LLMs with Extreme Sparsity
    2024/06/05 by Wentao Guo, Jikai Long, Guo, Wentao +21 · 4 citations
    Engineering · Mathematics · #Electromagnetic Simulation and Numerical Methods #Particle accelerators and beam dynamics #Numerical methods for differential equations
  31. JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
    2023/10/01 by Yuandong Tian, Yiping Wang, Tian, Yuandong +7 · 3 citations
    Computer Science · Engineering · #Advanced Memory and Neural Computing #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG)
  32. Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
    2024/04/01 by Harry Dong, Dong, Harry, Beidi Chen +3 · 3 citations
    Computer Science · Engineering · #Speech Recognition and Synthesis #Flow Measurement and Analysis #Neural Networks and Applications
  33. InRank: Incremental Low-Rank Learning
    2023/06/20 by Jiawei Zhao, Yifei Zhang, Zhao, Jiawei +7 · 2 citations
    Computer Science · Engineering · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and ELM #Sparse and Compressive Sensing Techniques
  34. The Last Human-Written Paper: Agent-Native Research Artifacts
    2026/04/27 by Jiachen Liu, Jiaxin Pei, Jintao Huang +34 · 3 voices
    Computer Science · #cs.LG
  35. Kinetics: Rethinking Test-Time Scaling Laws
    2025/06/05 by Ranajoy Sadhukhan, Sadhukhan, Ranajoy, Zhuoming Chen +9 · 6 citations
    Computer Science · Materials Science · #Cloud Computing and Resource Management #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Parallel Computing and Optimization Techniques
  36. S2FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
    2024/12/09 by Xinyu Yang, Jixuan Leng, Yang, Xinyu +12 · 2 citations
    Engineering · #Particle accelerators and beam dynamics #Electromagnetic Simulation and Numerical Methods #Particle Accelerators and Free-Electron Lasers
  37. Fast Algorithms for a New Relaxation of Optimal Transport
    2023/07/14 by Moses Charikar, Charikar, Moses, Beidi Chen +5 · 1 citation
    Computer Science · Mathematics · #Data Management and Algorithms #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #Markov Chains and Monte Carlo Methods #Stochastic processes and statistical mechanics
  38. It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
    2024/06/12 by Taiming Lu, Lingfeng Shen, Lu, Taiming +9 · 1 citation
    Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Healthcare Policy and Management #Machine Learning (cs.LG)
  39. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
    2026/07/20 by Krish Agarwal, Zhuoming Chen, Yanyuan Qin +3
    #cs.LG