Lu, Jiasen
- VQA: Visual Question Answering
2015/05/03 by Aishwarya Agrawal, Agrawal, Aishwarya, Jiasen Lu +11 · 378 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
2025/10/22 by Yusu Qian, Eli Bocek-Rivele, Qian, Yusu +13 · 5 voices · 8 citations
#cs.CV #cs.CL #cs.LG
- ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
2019/08/06 by Jiasen Lu, Dhruv Batra, Lu, Jiasen +5 · 135 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
2023/12/28 by Jiasen Lu, Lu, Jiasen, Christopher Clark +14 · 2 voices · 62 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition #Multimodal Machine Learning Applications #cs.AI #cs.CL #cs.CV
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
2024/09/25 by Matt Deitke, Christopher Clark, Deitke, Matt +100 · 4 voices · 151 citations
Computer Science · #Semantic Web and Ontologies #Speech and dialogue systems
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
2022/06/17 by Lu, Jiasen, Clark, Christopher, Zellers, Rowan +2 · 30 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Graph R-CNN for Scene Graph Generation
2018/08/01 by Jianwei Yang, Yang, Jianwei, Jiasen Lu +7 · 15 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Neural Network Applications #Advanced Image and Video Retrieval Techniques
- Hierarchical Question-Image Co-Attention for Visual Question Answering
2016/05/31 by Lu, Jiasen, Yang, Jianwei, Batra, Dhruv +1 · 12 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- One Diffusion to Generate Them All
2024/11/25 by Duong H. Le, Tuan Pham, Le, Duong H. +13 · 1 voice · 16 citations
#cs.CV #cs.AI
- MERLOT Reserve: Neural Script Knowledge through Vision and Language and\n Sound
2022/01/07 by Rowan Zellers, Jiasen Lu, Zellers, Rowan +17 · 15 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
- The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
2024/11/07 by Zhaofeng Wu, Xinyan Yu, Xinyan Velocity Yu +8 · 4 voices · 11 citations
Computer Science · #Natural Language Processing Techniques #cs.CL
- Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
2016/12/06 by Jiasen Lu, Lu, Jiasen, Caiming Xiong +5 · 15 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Image and Video Retrieval Techniques
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
2019/01/10 by Ma, Chih-Yao, Lu, Jiasen, Wu, Zuxuan +4 · 7 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
- Neural Baby Talk
2018/03/27 by Lu, Jiasen, Yang, Jianwei, Batra, Dhruv +1 · 6 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- ParlAI: A Dialog Research Software Platform
2017/05/18 by Miller, Alexander H., Feng, Will, Fisch, Adam +5 · 4 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- 12-in-1: Multi-Task Vision and Language Representation Learning
2019/12/05 by Jiasen Lu, Lu, Jiasen, Vedanuj Goswami +7 · 4 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
- A Simple Long-Tailed Recognition Baseline via Vision-Language Model
2021/11/29 by Teli Ma, Ma, Teli, Shijie Geng +13 · 4 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- Multi-Modal Answer Validation for Knowledge-Based VQA
2021/03/23 by Jialin Wu, Wu, Jialin, Jiasen Lu +5 · 3 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
- MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
2024/10/09 by Haihong Ye, Ye, Hanrong, Haotian Zhang +19 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
2017/06/05 by Jiasen Lu, Lu, Jiasen, Anitha Kannan +7 · 5 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Artificial Intelligence in Games
- SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
2025/03/24 by Xu, Mingze, Gao, Mingfei, Li, Shiyu +7 · 9 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- STIV: Scalable Text and Image Conditioned Video Generation
2024/12/10 by Lin, Zongyu, Liu, Wei, Chen, Chen +13 · 5 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
- Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition
2018/10/01 by Jianwei Yang, Yang, Jianwei, Jiasen Lu +7 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Robotics (cs.RO)
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
2025/05/20 by Mingfei Gao, Tian, Rui, Mingze Xu +12 · 6 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems
- Spatially Aware Multimodal Transformers for TextVQA
2020/07/23 by Kant, Yash, Batra, Dhruv, Anderson, Peter +4 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
2020/09/23 by Jaemin Cho, Cho, Jaemin, Jiasen Lu +7 · 2 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- AToken: A Unified Tokenizer for Vision
2025/09/17 by Lu, Jiasen, Song, Liangchen, Xu, Mingze +5 · 8 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
- GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
2025/05/16 by Yusu Qian, Jiasen Lü, Qian, Yusu +12 · 2 citations
Arts and Humanities · Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
2025/09/23 by Chen Chen, Chen, Chen, Pengsheng Guo +17 · 1 citation
Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning