vix.ing · top · new · best · stats · spec

Chuofan Ma

  1. Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
    2024/04/19 by Chuofan Ma, Ma, Chuofan, Yi Jiang +7 · 37 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  2. UniTok: A Unified Tokenizer for Visual Generation and Understanding
    2025/02/27 by Chuofan Ma, Yi Jiang, Ma, Chuofan +13 · 48 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  3. Liquid: Language Models are Scalable and Unified Multi-modal Generators
    2024/12/05 by Junfeng Wu, Yi Jiang, Wu, Junfeng +13 · 33 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
  4. CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection
    2023/10/25 by Chuofan Ma, Yi Jiang, Ma, Chuofan +7 · 1 voice · 6 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CV
  5. EGC: Image Generation and Classification via a Diffusion Energy-Based Model
    2023/04/04 by Qiushan Guo, Chuofan Ma, Guo, Qiushan +9 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection