vix.ing · top · new · best · stats · spec

QTrack: Query-Driven Reasoning for Multi-modal MOT

2026/03/14 by Tajamul Ashraf, Tavaheed Tariq, Sonia Yadav +4 · 1 voice
Computer Science · #Benchmark (surveying) #Coherence (philosophical gambling strategy) #Construct (python library) #Human Pose and Action Recognition #Identity (music) #Multimodal Machine Learning Applications #Natural language #Semantics (computer science) #Tracking (education) #Video Surveillance and Tracking Methods #cs.CV

paper · pdf · doi:10.48550/arxiv.2603.13759

openalex publication_date 2026/03/14 · arxiv published 2026/03/14 · arxiv updated 2026/03/14 · openalex created_date 2026/03/18 · openalex updated_date 2026/07/28

Abstract

Multi-object tracking (MOT) has traditionally focused on estimating trajectories of all objects in a video, without selectively reasoning about user-specified targets under semantic instructions. In this work, we introduce a query-driven tracking paradigm that formulates tracking as a spatiotemporal reasoning problem conditioned on natural language queries. Given a reference frame, a video sequence, and a textual query, the goal is to localize and track only the target(s) specified in the query while maintaining temporal coherence and identity consistency. To support this setting, we construct RMOT26, a large-scale benchmark with grounded queries and sequence-level splits to prevent identity leakage and enable robust evaluation of generalization. We further present QTrack, an end-to-end vision-language model that integrates multimodal reasoning with tracking-oriented localization. Additionally, we introduce a Temporal Perception-Aware Policy Optimization strategy with structured rewards to encourage motion-aware reasoning. Extensive experiments demonstrate the effectiveness of our approach for reasoning-centric, language-guided tracking. Code and data are available at https://github.com/gaash-lab/QTrack

Citations

Discussions

Related