vix.ing · top · new · best · stats · spec

On the use of human reference data for evaluating automatic image\n descriptions

2020/06/15 by Emiel van Miltenburg, van Miltenburg, Emiel
Computer Science · Medicine · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #COVID-19 diagnosis using AI

paper · pdf · doi:10.48550/arxiv.2006.08792

Abstract

Automatic image description systems are commonly trained and evaluated using\ncrowdsourced, human-generated image descriptions. The best-performing system is\nthen determined using some measure of similarity to the reference data (BLEU,\nMeteor, CIDER, etc). Thus, both the quality of the systems as well as the\nquality of the evaluation depends on the quality of the descriptions. As\nSection 2 will show, the quality of current image description datasets is\ninsufficient. I argue that there is a need for more detailed guidelines that\ntake into account the needs of visually impaired users, but also the\nfeasibility of generating suitable descriptions. With high-quality data,\nevaluation of image description systems could use reference descriptions, but\nwe should also look for alternatives.\n

Related