vix.ing · top · new · best · stats · spec

Shimao Chen

  1. MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
    2025/05/12 by LLM-Core Xiaomi, Xiaomi, LLM-Core, : +118 · 1 voice · 32 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
  2. MiMo-VL Technical Report
    2025/06/04 by Zihao Yue, Core Team, Zhenru Lin +135 · 18 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Reinforcement Learning in Robotics
  3. iFairy: the First 2-bit Complex LLM with All Parameters in \±1, ± i\
    2025/08/07 by Wang, Feiyu, Guoan Wang, Wang, Guoan +14 · 1 voice · 3 citations
    Computer Science · Physics and Astronomy · #Advanced Data Storage Technologies #Algorithms and Data Compression #Magnetic confinement fusion research
  4. INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
    2024/09/25 by Shimao Chen, Chen, Shimao, Zirui Liu +18 · 1 voice · 1 citation
    Computer Science · Engineering · Neuroscience · #Brain Tumor Detection and Classification #Image Processing Techniques and Applications #Image and Signal Denoising Methods #cs.AI #cs.LG
  5. Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
    2024/04/25 by Peizhuang Cong, Aomufei Yuan, Cong, Peizhuang +9 · 3 citations
    Business, Management and Accounting · Decision Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computation and Language (cs.CL) #FOS: Computer and information sciences #Forecasting Techniques and Applications #Machine Learning (cs.LG)
  6. HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
    2026/02/03 by Yizhao Gao, Jianyu Wei, Qihao Zhang +11 · 1 voice
    #cs.CL #cs.AI