2026/01/03 by Hossein Amiri, Joon-Seok Kim, Hamdi Kavak +4
#cs.SE
Understanding individual-level human mobility is critical for a wide range of applications. Real-world trajectory datasets provide valuable insights into movement behaviors and patterns of life but are often constrained by data sparsity and participation bias. Synthetic data, by contrast, offers scalability and flexibility but frequently lacks realism. % To address this gap, we introduce a comprehensive software pipeline for generating, calibrating, processing, and visualizing large-scale individual-level human mobility datasets that combine the realism of empirical data with the control and extensibility simulations. % Our system consists of four integrated components: (1) a data generation engine that constructs geographically grounded simulations using OpenStreetMap data to produce diverse mobility logs; (2) a genetic algorithm--based calibration module that fine-tunes simulation parameters to align with real-world mobility characteristics; (3) a data processing suite that transforms raw simulation logs into structured formats suitable for downstream applications; and (4) a visualization module that extracts and presents key mobility patterns and insights from the processed datasets for improved interpretability. Evaluation of generated trajectory datasets for the Atlanta, Georgia, USA region show realistic behavior that, despite emerging from a simulation without any reference to real human individuals, exhibits realistic human behavior that closely matches aggregate metrics of real-world datasets. We also provide a sensitivity analysis to study what simulation parameters affect simulation runtime. Code and simulated datasets are shared to provide the broad research community with large-scale dataset that, albeit not real, exhibit realistic human behavior while being orders of magnitudes larger than any open real-world mobility dataset.