2020/05/14 by Kristian Muri Knausgård, Knausgård, Kristian Muri, Arne Wiklund +11 · 1 citation
Biochemistry, Genetics and Molecular Biology · Environmental Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Ichthyology and Marine Biology #Identification and Quantification in Food #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #Water Quality Monitoring Technologies #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2005.07518
openalex publication_date 2020/05/14 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
A wide range of applications in marine ecology extensively uses underwater\ncameras. Still, to efficiently process the vast amount of data generated, we\nneed to develop tools that can automatically detect and recognize species\ncaptured on film. Classifying fish species from videos and images in natural\nenvironments can be challenging because of noise and variation in illumination\nand the surrounding habitat. In this paper, we propose a two-step deep learning\napproach for the detection and classification of temperate fishes without\npre-filtering. The first step is to detect each single fish in an image,\nindependent of species and sex. For this purpose, we employ the You Only Look\nOnce (YOLO) object detection technique. In the second step, we adopt a\nConvolutional Neural Network (CNN) with the Squeeze-and-Excitation (SE)\narchitecture for classifying each fish in the image without pre-filtering. We\napply transfer learning to overcome the limited training samples of temperate\nfishes and to improve the accuracy of the classification. This is done by\ntraining the object detection model with ImageNet and the fish classifier via a\npublic dataset (Fish4Knowledge), whereupon both the object detection and\nclassifier are updated with temperate fishes of interest. The weights obtained\nfrom pre-training are applied to post-training as a priori. Our solution\nachieves the state-of-the-art accuracy of 99.27 % on the pre-training. The\npercentage values for accuracy on the post-training are good; 83.68 % and\n87.74 % with and without image augmentation, respectively, indicating that the\nsolution is viable with a more extensive dataset.\n