ByteTrack: Multi-Object Tracking by Associating Every Detection Box
2021/10/13 by Yifu Zhang, Zhang, Yifu, Peize Sun +15 · 165 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #UAV Applications and Optimization #Video Surveillance and Tracking Methods #cs.CV
paper · pdf · doi:10.48550/arxiv.2110.06864
openalex publication_date 2021/10/13 · arxiv created 2022/04/07 · arxiv updated 2022/04/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Multi-object tracking (MOT) aims at estimating bounding boxes and identities of objects in videos. Most methods obtain identities by associating detection boxes whose scores are higher than a threshold. The objects with low detection scores, e.g. occluded objects, are simply thrown away, which brings non-negligible true object missing and fragmented trajectories. To solve this problem, we present a simple, effective and generic association method, tracking by associating almost every detection box instead of only the high score ones. For the low score detection boxes, we utilize their similarities with tracklets to recover true objects and filter out the background detections. When applied to 9 different state-of-the-art trackers, our method achieves consistent improvement on IDF1 score ranging from 1 to 10 points. To put forwards the state-of-the-art performance of MOT, we design a simple and strong tracker, named ByteTrack. For the first time, we achieve 80.3 MOTA, 77.3 IDF1 and 63.1 HOTA on the test set of MOT17 with 30 FPS running speed on a single V100 GPU. ByteTrack also achieves state-of-the-art performance on MOT20, HiEve and BDD100K tracking benchmarks. The source code, pre-trained models with deploy versions and tutorials of applying to other trackers are released at https://github.com/ifzhang/ByteTrack.
Citations
Cited by
- Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking
- Learning Association via Track-Detection Matching for Multi-Object Tracking
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- Training-Free Global Geometric Association for 4D LiDAR Panoptic Segmentation
- Object-Centric Framework for Video Moment Retrieval
- ST-DETrack: Identity-Preserving Branch Tracking in Entangled Plant Canopies via Dual Spatiotemporal Evidence
- TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Team-Aware Football Player Tracking with SAM: An Appearance-Based Approach to Occlusion Recovery
- GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring
- DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- SDG-Track: A Heterogeneous Observer-Follower Framework for High-Resolution UAV Tracking on Embedded Platforms
- EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
- From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Physical ID-Transfer Attacks against Multi-Object Tracking via Adversarial Trajectory
- TransientTrack: Advanced Multi-Object Tracking and Classification of Cancer Cells with Transient Fluorescent Signals
- MANTA: Physics-Informed Generalized Underwater Object Tracking
- ByteStorm: a multi-step data-driven approach for Tropical Cyclones detection and tracking
- DM3T: Harmonizing Modalities via Diffusion for Multi-Object Tracking
- StableTrack: Stabilizing Multi-Object Tracking on Low-Frequency Detections
- SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
- Occlusion-Aware Multi-Object Tracking via Expected Probability of Detection
- DynaMix: Generalizable Person Re-identification via Dynamic Relabeling and Mixed Data Sampling
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Stable Multi-Drone GNSS Tracking System for Marine Robots
- A Tri-Modal Dataset and a Baseline System for Tracking Unmanned Aerial Vehicles
- Tracking and Segmenting Anything in Any Modality
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
- Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models
- Enhancing Multi-Camera Gymnast Tracking Through Domain Knowledge Integration
- YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras
- StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
- The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification
- Deep Learning for Accurate Vision-based Catch Composition in Tropical Tuna Purse Seiners
- SVBRD-LLM: Self-Verifying Behavioral Rule Discovery for Autonomous Vehicle Identification
- Visionary Co-Driver: Enhancing Driver Perception of Potential Risks with LLM and HUD
- Edge Assisted Multi-Camera Vehicle Tracking Framework for Real-Time and Scalable Deployment
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- Ground Plane Projection for Improved Traffic Analytics at Intersections
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery
- DMSORT: An efficient parallel maritime multi-object tracking architecture for unmanned vessel platforms
- Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
- Zero-Shot Multi-Animal Tracking in the Wild
- OmniTrack++: Omnidirectional Multi-Object Tracking by Learning Large-FoV Trajectory Feedback
- IoT- and AI-informed urban air quality models for vehicle pollution monitoring
- When Fish Look Alike: Tracking Identities with Dual-branch Elasticity
- WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps
- Cataract-LMM: Large-Scale, Multi-Source, Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
- A Hybrid Approach for Visual Multi-Object Tracking
- GenTrack: A New Generation of Multi-Object Tracking
- Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
- GRAP-MOT: Unsupervised Graph-based Position Weighted Person Multi-camera Multi-object Tracking in a Highly Congested Space
- Physics-Guided Fusion for Robust 3D Tracking of Fast Moving Small Objects
- Multi-Camera Worker Tracking in Logistics Warehouse Considering Wide-Angle Distortion
- Symmetric Entropy-Constrained Video Coding for Machines
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- DMTrack: Deformable State-Space Modeling for UAV Multi-Object Tracking with Kalman Fusion and Uncertainty-Aware Association
- EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
- MMOT: The First Challenging Benchmark for Drone-based Multispectral Multi-Object Tracking
- Fast Self-Supervised depth and mask aware Association for Multi-Object Tracking
- GL-DT: Multi-UAV Detection and Tracking with Global-Local Integration
- Cattle-CLIP: A Multimodal Framework for Cattle Behaviour Recognition from Video
- Human Action Recognition from Point Clouds over Time
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
- A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
- Motion-Aware Transformer for Multi-Object Tracking
- Design Insights and Comparative Evaluation of a Hardware-Based Cooperative Perception Architecture for Lane Change Prediction
- iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
- DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
- An Analysis of Kalman Filter based Object Tracking Methods for Fast-Moving Tiny Objects
- Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- VSE-MOT: Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Enhancement
- Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation
- Motion Estimation for Multi-Object Tracking using KalmanNet with Semantic-Independent Encoding
- Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
- An HMM-based framework for identity-aware long-term multi-object tracking from sparse and uncertain identification: use case on long-term tracking in livestock
- UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
- NOOUGAT: Towards Unified Online and Offline Multi-Object Tracking
- FLUID: A Fine-Grained Lightweight Urban Signalized-Intersection Dataset of Dense Conflict Trajectories
- Contrastive Learning through Auxiliary Branch for Video Object Detection
- SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
- FastTracker: Real-Time and Accurate Visual Tracking
- SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
- SocialTrack: Multi-Object Tracking in Complex Urban Traffic Scenes Inspired by Social Behavior
- Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
- VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
- MeMoSORT: Memory-Assisted Filtering and Motion-Adaptive Association Metric for Multi-Person Tracking
- GRASPTrack: Geometry-Reasoned Association via Segmentation and Projection for Multi-Object Tracking
- TrackOR: Towards Personalized Intelligent Operating Rooms Through Robust Tracking
- Head Anchor Enhanced Detection and Association for Crowded Pedestrian Tracking
- Multi-tracklet Tracking for Generic Targets with Adaptive Detection Clustering
- Benchmarking pig detection and tracking under diverse and challenging conditions
- Drone Detection with Event Cameras
- Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos
- DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion
- Stable at Any Speed: Speed-Driven Multi-Object Tracking with Learnable Kalman Filtering
- Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
- DRWKV: Focusing on Object Edges for Low-Light Image Enhancement
- Real-Time Fusion of Visual and Chart Data for Enhanced Maritime Vision
- Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations
- MVA 2025 Small Multi-Object Tracking for Spotting Birds Challenge: Dataset, Methods, and Results
- YOLOv8-SMOT: An Efficient and Robust Framework for Real-Time Small Object Tracking via Slice-Assisted Training and Adaptive Association
- Predicting Soccer Penalty Kick Direction Using Human Action Recognition
- Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision Users
- Glance-MCMT: A General MCMT Framework with Glance Initialization and Progressive Association
- Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder
- Unified People Tracking with Graph Neural Networks
- RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking
- Towards Continuous Home Cage Monitoring: An Evaluation of Tracking and Identification Strategies for Laboratory Mice
- When Trackers Date Fish: A Benchmark and Framework for Underwater Multiple Fish Tracking
- CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
- A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
- Autoregressive Denoising Score Matching is a Good Video Anomaly Detector
- SAM2Auto: Auto Annotation Using FLASH
- Automating Traffic Monitoring with SHM Sensor Networks via Vision-Supervised Deep Learning
- USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
- NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- Open-World Object Counting in Videos
- Towards Perception-based Collision Avoidance for UAVs when Guiding the Visually Impaired
- Deep Learning-Based Multi-Object Tracking: A Comprehensive Survey from Foundations to State-of-the-Art
- Video Individual Counting With Implicit One-to-Many Matching
- Real-World Deployment of a Lane Change Prediction Architecture Based on Knowledge Graph Embeddings and Bayesian Inference
- Multiple Object Tracking in Video SAR: A Benchmark and Tracking Baseline
- FRED: The Florence RGB-Event Drone Dataset
- SMMT: Siamese Motion Mamba with Self-attention for Thermal Infrared Target Tracking
- Person Recognition at Altitude and Range: Fusion of Face, Body Shape and Gait
- High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset
- SportMamba: Adaptive Non-Linear Multi-Object Tracking with State Space Models for Team Sports
- No Train Yet Gain: Towards Generic Multi-Object Tracking in Sports and Beyond
- ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised Detection
- Towards Edge-Based Idle State Detection in Construction Machinery Using Surveillance Cameras
- Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking
- GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
- ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking
- Underwater moving target detection using online robust principal component analysis and multimodal anomaly detection
- Using Cross-Domain Detection Loss to Infer Multi-Scale Information for Improved Tiny Head Tracking
- TPT-Bench: A Large-Scale, Long-Term and Robot-Egocentric Dataset for Benchmarking Target Person Tracking
- VALISENS: A Validated Innovative Multi-Sensor System for Cooperative Automated Driving
- PaniCar: Securing the Perception of Advanced Driving Assistance Systems Against Emergency Vehicle Lighting
- CAMELTrack: Context-Aware Multi-cue ExpLoitation for Online Multi-Object Tracking
- Experimental study on surveillance video-based indoor occupancy measurement with occupant-centric control
- Visual Trajectory Prediction of Vessels for Inland Navigation
- Vehicular Communication Security: Multi-Channel and Multi-Factor Authentication
- Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
- A computer vision method to estimate ventilation rate of Atlantic salmon in sea fish farms
- PerfCam: Digital Twinning for Production Lines Using 3D Gaussian Splatting and Vision Models
- Designing Digital Humans with Ambient Intelligence
- EffOWT: Transfer Visual Language Models to Open-World Tracking Efficiently and Effectively
- SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation
- Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions
- DRIFT open dataset: A drone-derived intelligence for traffic analysis in urban environment
- WildLive: Near Real-time Visual Wildlife Tracking onboard UAVs
- PapMOT: Exploring Adversarial Patch Attack against Multiple Object Tracking
- Automatic recognition of isolated piglet outliers based on multi-object tracking
Related