2016/10/25 by Suchet Bargoti, Bargoti, Suchet, James Underwood +1 · 1 citation
Agricultural and Biological Sciences · Chemistry · #Computer Vision and Pattern Recognition (cs.CV) #Date Palm Research Studies #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO) #Smart Agriculture and AI #Spectroscopy and Chemometric Analyses
paper · pdf · doi:10.48550/arxiv.1610.08120
openalex publication_date 2016/10/25 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28
Ground vehicles equipped with monocular vision systems are a valuable source\nof high resolution image data for precision agriculture applications in\norchards. This paper presents an image processing framework for fruit detection\nand counting using orchard image data. A general purpose image segmentation\napproach is used, including two feature learning algorithms; multi-scale\nMulti-Layered Perceptrons (MLP) and Convolutional Neural Networks (CNN). These\nnetworks were extended by including contextual information about how the image\ndata was captured (metadata), which correlates with some of the appearance\nvariations and/or class distributions observed in the data. The pixel-wise\nfruit segmentation output is processed using the Watershed Segmentation (WS)\nand Circular Hough Transform (CHT) algorithms to detect and count individual\nfruits. Experiments were conducted in a commercial apple orchard near\nMelbourne, Australia. The results show an improvement in fruit segmentation\nperformance with the inclusion of metadata on the previously benchmarked MLP\nnetwork. We extend this work with CNNs, bringing agrovision closer to the\nstate-of-the-art in computer vision, where although metadata had negligible\ninfluence, the best pixel-wise F1-score of 0.791 was achieved. The WS\nalgorithm produced the best apple detection and counting results, with a\ndetection F1-score of 0.858. As a final step, image fruit counts were\naccumulated over multiple rows at the orchard and compared against the\npost-harvest fruit counts that were obtained from a grading and counting\nmachine. The count estimates using CNN and WS resulted in the best performance\nfor this dataset, with a squared correlation coefficient of r2=0.826.\n