2018/12/16 by Ehsan Abbasnejad, Iman Abbasnejad, Abbasnejad, Ehsan +7 · 1 citation
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1812.06398
openalex publication_date 2018/12/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
As Computer Vision moves from a passive analysis of pixels to active analysis\nof semantics, the breadth of information algorithms need to reason over has\nexpanded significantly. One of the key challenges in this vein is the ability\nto identify the information required to make a decision, and select an action\nthat will recover it. We propose a reinforcement-learning approach that\nmaintains a distribution over its internal information, thus explicitly\nrepresenting the ambiguity in what it knows, and needs to know, towards\nachieving its goal. Potential actions are then generated according to this\ndistribution. For each potential action a distribution of the expected outcomes\nis calculated, and the value of the potential information gain assessed. The\naction taken is that which maximizes the potential information gain. We\ndemonstrate this approach applied to two vision-and-language problems that have\nattracted significant recent interest, visual dialog and visual query\ngeneration. In both cases, the method actively selects actions that will best\nreduce its internal uncertainty and outperforms its competitors in achieving\nthe goal of the challenge.\n