Detect Anything via Next Point Prediction
2025/10/14 by Qing Jiang, Jiang, Qing, Huo, Junan +14 · 10 citations
Computer Science · Neuroscience · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face Recognition and Perception #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2510.12798
openalex publication_date 2025/10/14 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28
Abstract
Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to leverage MLLMs to tackle this task, they face challenges like low recall rate, duplicate predictions, coordinate misalignment, etc. In this work, we bridge this gap and propose Rex-Omni, a 3B-scale MLLM that achieves state-of-the-art object perception performance. On benchmarks like COCO and LVIS, Rex-Omni attains performance comparable to or exceeding regression-based models (e.g., DINO, Grounding DINO) in a zero-shot setting. This is enabled by three key designs: 1) Task Formulation: we use special tokens to represent quantized coordinates from 0 to 999, reducing the model's learning difficulty and improving token efficiency for coordinate prediction; 2) Data Engines: we construct multiple data engines to generate high-quality grounding, referring, and pointing data, providing semantically rich supervision for training; \3) Training Pipelines: we employ a two-stage training process, combining supervised fine-tuning on 22 million data with GRPO-based reinforcement post-training. This RL post-training leverages geometry-aware rewards to effectively bridge the discrete-to-continuous coordinate prediction gap, improve box accuracy, and mitigate undesirable behaviors like duplicate predictions that stem from the teacher-guided nature of the initial SFT stage. Beyond conventional detection, Rex-Omni's inherent language understanding enables versatile capabilities such as object referring, pointing, visual prompting, GUI grounding, spatial referring, OCR and key-pointing, all systematically evaluated on dedicated benchmarks. We believe that Rex-Omni paves the way for more versatile and language-aware visual perception systems.
Citations
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Ovis2.5 Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- PaddleOCR 3.0 Technical Report
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
- Seed1.5-VL Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- Qwen2.5-VL Technical Report
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
- DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
- YOLOv11: An Overview of the Key Architectural Enhancements
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
- Aria: An Open Multimodal Native Mixture-of-Experts Model
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
- SAM 2: Segment Anything in Images and Videos
- RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
- Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
- DAVE -- A Detect-and-Verify Paradigm for Low-Shot Counting
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
- Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
- T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- YOLO-World: Real-Time Open-Vocabulary Object Detection
- APTv2: Benchmarking Animal Pose Estimation and Tracking with a Large-scale Dataset and Beyond
- Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
- T-Rex: Counting by Visual Prompting
- CogVLM: Visual Expert for Pretrained Language Models
- CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding
- LISA: Reasoning Segmentation via Large Language Model
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- Scaling Open-Vocabulary Object Detection
- M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis
- DETRs Beat YOLOs on Real-time Object Detection
- DETRs Beat YOLOs on Real-time Object Detection
- Segment Everything Everywhere All at Once
- V3Det: Vast Vocabulary Visual Detection Dataset
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes
- PACO: Parts and Attributes of Common Objects
- DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Few-shot Object Counting and Detection
- APT-36K: A Large-scale Benchmark for Animal Pose Estimation and Tracking
- Towards End-to-End Unified Scene Text Detection and Layout Analysis
- Learning Affordance Grounding from Exocentric Images
- Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
- OCR-IDL: OCR Annotations for Industry Document Library Dataset
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
- Scaling Open-Vocabulary Image Segmentation with Image-Level Labels
- RegionCLIP: Region-based Language-Image Pretraining
- Grounded Language-Image Pre-training
- PartImageNet: A Large, High-Quality Dataset of Parts
- Pix2seq: A Language Modeling Framework for Object Detection
- AP-10K: A Benchmark for Animal Pose Estimation in the Wild
- UIBert: Learning Generic Multimodal Representations for UI Understanding
- Dynamic Head: Unifying Object Detection Heads with Attentions
- TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped\n scene text
- MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding
- Learning To Count Everything
- Spatial Dual-Modality Graph Reasoning for Key Information Extraction
- FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery
- Learning Transferable Visual Models From Natural Language Supervision
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- End-to-End Object Detection with Transformers
- EfficientDet: Scalable and Efficient Object Detection
- Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning
- ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text (RRC-ArT)
- ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
- PubLayNet: largest dataset ever for document layout analysis
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- CenterNet: Keypoint Triplets for Object Detection
- FCOS: Fully Convolutional One-Stage Object Detection
- nuScenes: A multimodal dataset for autonomous driving
- TableBank: A Benchmark Dataset for Table Detection and Recognition
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- Detector-in-Detector: Multi-Level Analysis for Human-Parts
- CornerNet: Detecting Objects as Paired Keypoints
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- CrowdHuman: A Benchmark for Detecting Human in a Crowd
- YOLOv3: An Incremental Improvement
- mixup: Beyond Empirical Risk Minimization
- ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)
- Feature Pyramid Networks for Object Detection
- Generation and Comprehension of Unambiguous Object Descriptions
- You Only Look Once: Unified, Real-Time Object Detection
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Fast R-CNN
- Microsoft COCO: Common Objects in Context
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related