2021/11/16 by Samayan Bhattacharya, Bhattacharya, Samayan, Sk Shahnawaz +1
Computer Science · Engineering · Mathematics · #Artificial intelligence #Bottleneck #Cluster analysis #Computer science #Computer vision #Domain Adaptation and Few-Shot Learning #Exploit #Hand Gesture Recognition Systems #Human Pose and Action Recognition #Machine learning #Mathematics #Pattern recognition (psychology) #Pose #Robotics and Sensor-Based Localization #Unavailability #cs.AI #cs.CV
paper · pdf · doi:10.48550/arxiv.2111.08259
published in arXiv (Cornell University) (Cornell University) · 9 pages, 4 figures, 3 tables
arxiv created 2021/11/16 · openalex publication_date 2021/11/16 · arxiv updated 2021/11/17 · openalex created_date 2022/10/22 · openalex updated_date 2026/08/08
Animal pose estimation has recently come into the limelight due to its\napplication in biology, zoology, and aquaculture. Deep learning methods have\neffectively been applied to human pose estimation. However, the major\nbottleneck to the application of these methods to animal pose estimation is the\nunavailability of sufficient quantities of labeled data. Though there are ample\nquantities of unlabelled data publicly available, it is economically\nimpractical to label large quantities of data for each animal. In addition, due\nto the wide variety of body shapes in the animal kingdom, the transfer of\nknowledge across domains is ineffective. Given the fact that the human brain is\nable to recognize animal pose without requiring large amounts of labeled data,\nit is only reasonable that we exploit unsupervised learning to tackle the\nproblem of animal pose recognition from the available, unlabelled data. In this\npaper, we introduce a novel architecture that is able to recognize the pose of\nmultiple animals fromunlabelled data. We do this by (1) removing background\ninformation from each image and employing an edge detection algorithm on the\nbody of the animal, (2) Tracking motion of the edge pixels and performing\nagglomerative clustering to segment body parts, (3) employing contrastive\nlearning to discourage grouping of distant body parts together. Hence we are\nable to distinguish between body parts of the animal, based on their visual\nbehavior, instead of the underlying anatomy. Thus, we are able to achieve a\nmore effective classification of the data than their human-labeled\ncounterparts. We test our model on the TigDog and WLD (WildLife Documentary)\ndatasets, where we outperform state-of-the-art approaches by a significant\nmargin. We also study the performance of our model on other public data to\ndemonstrate the generalization ability of our model.\n