Chuofan Ma
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
2024/04/19 by Chuofan Ma, Ma, Chuofan, Yi Jiang +7 · 37 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
2025/02/27 by Chuofan Ma, Yi Jiang, Ma, Chuofan +13 · 48 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- Liquid: Language Models are Scalable and Unified Multi-modal Generators
2024/12/05 by Junfeng Wu, Yi Jiang, Wu, Junfeng +13 · 33 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
- CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection
2023/10/25 by Chuofan Ma, Yi Jiang, Ma, Chuofan +7 · 1 voice · 6 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CV
- EGC: Image Generation and Classification via a Diffusion Energy-Based Model
2023/04/04 by Qiushan Guo, Chuofan Ma, Guo, Qiushan +9 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection