2021/10/24 by Md Aminul Haque Palash, Palash, Md Aminul Haque, MD Abdullah Al Nasim +9
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #cs.AI #cs.CV
paper · pdf · doi:10.48550/arxiv.2110.12442
15 pages, 6 figures, 1 table, 6 equations
arxiv created 2021/10/24 · openalex publication_date 2021/10/24 · arxiv updated 2021/10/26 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Automatic Image Captioning is the never-ending effort of creating syntactically and validating the accuracy of textual descriptions of an image in natural language with context. The encoder-decoder structure used throughout existing Bengali Image Captioning (BIC) research utilized abstract image feature vectors as the encoder's input. We propose a novel transformer-based architecture with an attention mechanism with a pre-trained ResNet-101 model image encoder for feature extraction from images. Experiments demonstrate that the language decoder in our technique captures fine-grained information in the caption and, then paired with image features, produces accurate and diverse captions on the BanglaLekhaImageCaptions dataset. Our approach outperforms all existing Bengali Image Captioning work and sets a new benchmark by scoring 0.694 on BLEU-1, 0.630 on BLEU-2, 0.582 on BLEU-3, and 0.337 on METEOR.