vix.ing · top · new · best · stats · spec

Shufan Li

  1. Mercury: Ultra-Fast Language Models Based on Diffusion
    2025/06/17 by Inception Labs, Samar Khanna, Labs, Inception +23 · 26 voices · 89 citations
    #cs.CL #cs.AI #cs.LG
  2. Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning
    2022/12/30 by Colorado Reed, Ritwik Gupta, Reed, Colorado J. +17 · 64 citations
    Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Synthetic Aperture Radar (SAR) Applications and Techniques
  3. Aligning Diffusion Models by Optimizing Human Utility
    2024/04/06 by Shufan Li, Konstantinos Kallidromitis, Li, Shufan +7 · 22 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cell Image Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  4. Hierarchical Open-vocabulary Universal Image Segmentation
    2023/07/03 by Xudong Wang, Wang, Xudong, Shufan Li +9 · 14 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  5. Mamba-ND: Selective State Space Modeling for Multi-Dimensional Data
    2024/02/08 by Shufan Li, Li, Shufan, Harkanwar Singh +3 · 17 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Neural Networks and Applications
  6. LaViDa: A Large Diffusion Language Model for Multimodal Understanding
    2025/05/22 by Shufan Li, Li, Shufan, Konstantinos Kallidromitis +17 · 41 citations
    #cs.CV
  7. SegLLM: Multi-round Reasoning Segmentation
    2024/10/24 by Xudong Wang, Wang, XuDong, Shaolun Zhang +13 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Fuzzy Logic and Control Systems #Rough Sets and Fuzzy Logic #Semantic Web and Ontologies
  8. InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
    2023/12/11 by Shufan Li, Li, Shufan, Harkanwar Singh +3 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  9. MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
    2024/12/17 by Hritik Bansal, Bansal, Hritik, Daniel Israel +9 · 2 citations
    Health Professions · #Artificial Intelligence (cs.AI) #Assistive Technology in Communication and Mobility #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  10. xT: Nested Tokenization for Larger Context in Large Images
    2024/03/04 by Ritwik Gupta, Shufan Li, Gupta, Ritwik +9 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
  11. Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
    2025/09/23 by Shufan Li, Li, Shufan, Jiuxiang Gu +11 · 5 citations
    #cs.CV
  12. Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
    2025/12/16 by Shufan Li, Li, Shufan, Jiuxiang Gu +11 · 2 citations
    #cs.CV
  13. Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory
    2024/11/25 by Eric Hanchen Jiang, Jiang, Eric Hanchen, Yasi Zhang +11 · 1 citation
    Computer Science · #Video Analysis and Summarization #Image Retrieval and Classification Techniques #Natural Language Processing Techniques
  14. PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction
    2025/06/18 by Shufan Li, Li, Shufan, Aditya Grover +1 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering