vix.ing · top · new · best · stats · spec

From Multimodal to Unimodal Webpages for Developing Countries

2017/11/06 by Sandeep Vidyapu, Vidyapu Sandeep, V Vijaya Saradhi +5
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Image Retrieval and Classification Techniques #Machine Learning (stat.ML) #Visual Attention and Saliency Detection #cs.HC #stat.ML

paper · pdf · doi:10.48550/arxiv.1711.02068

Presented at NIPS 2017 Workshop on Machine Learning for the Developing World

arxiv created 2017/11/06 · openalex publication_date 2017/11/06 · arxiv updated 2017/11/09 · openalex created_date 2017/11/17 · openalex updated_date 2026/07/28

Abstract

The multimodal web elements such as text and images are associated with inherent memory costs to store and transfer over the Internet. With the limited network connectivity in developing countries, webpage rendering gets delayed in the presence of high-memory demanding elements such as images (relative to text). To overcome this limitation, we propose a Canonical Correlation Analysis (CCA) based computational approach to replace high-cost modality with an equivalent low-cost modality. Our model learns a common subspace for low-cost and high-cost modalities that maximizes the correlation between their visual features. The obtained common subspace is used for determining the low-cost (text) element of a given high-cost (image) element for the replacement. We analyze the cost-saving performance of the proposed approach through an eye-tracking experiment conducted on real-world webpages. Our approach reduces the memory-cost by at least 83.35% by replacing images with text.

Citations

Related