2023/12/01 by Yutong Bai, Xinyang Geng, Bai, Yutong +14 · 1 voice · 28 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #cs.CV
paper · pdf · doi:10.48550/arxiv.2312.00785
openalex publication_date 2023/12/01 · arxiv published 2023/12/01 · arxiv updated 2023/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data. To do this, we define a common format, "visual sentences", in which we can represent raw images and videos as well as annotated data sources such as semantic segmentations and depth reconstructions without needing any meta-knowledge beyond the pixels. Once this wide variety of visual data (comprising 420 billion tokens) is represented as sequences, the model can be trained to minimize a cross-entropy loss for next token prediction. By training across various scales of model architecture and data diversity, we provide empirical evidence that our models scale effectively. Many different vision tasks can be solved by designing suitable visual prompts at test time.