vix.ing · top · new · best · stats · spec

M2FN: Multi-step modality fusion for advertisement image assessment

2021/01/25 by Kyung-Wha Park, Jung-Woo Ha, Junghoon Lee +5
Computer Science · Neuroscience · #Aesthetic Perception and Analysis #Artificial intelligence #Artificial neural network #Computer science #Data mining #Fusion #Image (mathematics) #Image Retrieval and Classification Techniques #Information retrieval #Machine learning #Modality (human–computer interaction) #Normalization (sociology) #Pairwise comparison #Pattern recognition (psychology) #Visual Attention and Saliency Detection #cs.AI #cs.CV

paper · pdf · doi:10.1016/j.asoc.2021.107116

published in Applied Soft Computing

openalex publication_date 2021/01/25 · arxiv created 2021/02/09 · arxiv updated 2021/02/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Assessing advertisements, specifically on the basis of user preferences and ad quality, is crucial to the marketing industry. Although recent studies have attempted to use deep neural networks for this purpose, these studies have not utilized image-related auxiliary attributes, which include embedded text frequently found in ad images. We, therefore, investigated the influence of these attributes on ad image preferences. First, we analyzed large-scale real-world ad log data and, based on our findings, proposed a novel multi-step modality fusion network (M2FN) that determines advertising images likely to appeal to user preferences. Our method utilizes auxiliary attributes through multiple steps in the network, which include conditional batch normalization-based low-level fusion and attention-based high-level fusion. We verified M2FN on the AVA dataset, which is widely used for aesthetic image assessment, and then demonstrated that M2FN can achieve state-of-the-art performance in preference prediction using a real-world ad dataset with rich auxiliary attributes.

Citations