2025/11/18 by Juho Korkeala, Korkeala, Juho, Jesse Muhojoki +11 · 1 citation
Agricultural and Biological Sciences · Environmental Science · #Convex hull #Laser scanning #Pattern recognition (psychology) #Plant Surface Properties and Treatments #Point cloud #Reference data #Remote Sensing and LiDAR Applications #Remote Sensing in Agriculture #Table (database) #Tree (set theory)
paper · pdf · doi:10.48550/arxiv.2512.05610
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/11/18 · openalex created_date 2025/12/09 · openalex updated_date 2026/08/05
General description This dataset consists of 1915 samples of tree point clouds across 8 species (pine, spruce, birch, maple, aspen, rowan, alder and lime). The data was collected from Espoonlahti, Espoo, Finland in June 2024. The measurement areas contain urban, semiurban and natural environments. The data was collected by a backpack-mounted laser scanner (FARO orbis, wavelength = 905 nm). The dataset was used in the publication: NormalView: sensor-agnostic tree species classification from backpack and aerial lidar data using geometric projections. Korkeala, J., Muhojoki, J., Taher, J., Salolahti, K., Hyyppä, M., Kukko, A. and Hyyppä, J. (2025). ArXiv preprint arXiv:2512.05610, doi: 10.48550/arXiv.2512.05610 Any scientific study using the data should cite the above article. Data The dataset consists of the following files: The tree segments are in the MLStreesegments.zip zip-file (1915 in total). The name of each .laz -file is the unique id used in the reference. The coordinates are centered so that the approximated tree trunk is in the origin. The reference data is in MLSspeciesref.csv -file. The reference file has 5 columns: id: The segment specific id of the sample. Species code: The species of the tree in the segment (see table below). height: Approximated height of the tree (evaluated as the difference between the 1st and 99th z-coordinate percentiles). xyarea: The 2D-area of the convex hull of the coordinates in the xy-plane. pointdensity: (Amount of points in a segment) / (xyarea of the segment). Count: Some trees were captured in multiple scans, and hence there are segments in the dataset which are different, but correspond to the same tree. The value of Count indicates how many times the tree in the segment is present in the dataset. If Count > 1, then there are other segments which correspond to the same tree. Most of the tree are captured only once. DBH: Diameter at breast height. Evaluated only for a part of the segments. The model weights for the models used in the study are in Modelweights.zip: The weights are stored as .pt (PyTorch) files, as given out by YOLOv11x by Ultralytics. The first part of the file name (ALS, ALS9, MLS, Imagesize) refers to the data the model was ttrained on. The second part (intensity1, normal, etc...) refers to the image creation method. Specieswisestatistics.pdf contains two species-level statistics plot derived from the tree segments. The mapping between species codes and actual species is Species code Species Sample count 1 Pine 755 2 Spruce 310 3 Birch 374 4 Maple 234 5 Aspen 108 6 Rowan 57 8 Lime 66 9 Alder 11 Note that the alders were not included in the classification tests in the paper due to low sample count. For further details on data acquisition and processing, we refer the reader to sections 2.1.-2.3. in Korkeala et. al. (arXiv:2512.05610, 2025). The tree segments are in LAS version 1.2 point format 3.