vix.ing · top · new · best · stats · spec

Tracing State-Level Obesity Prevalence from Sentence Embeddings of Tweets: A Feasibility Study

2019/11/26 by Xiaoyi Zhang, Zhang, Xiaoyi, Rodoniki Athanasiadou +3
Medicine · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Data-Driven Disease Surveillance #FOS: Computer and information sciences #Smoking Behavior and Cessation #Social and Information Networks (cs.SI)

paper · pdf · doi:10.48550/arxiv.1911.11324

openalex publication_date 2019/11/26 · openalex created_date 2019/12/05 · openalex updated_date 2026/07/28

Abstract

Twitter data has been shown broadly applicable for public health surveillance. Previous public health studies based on Twitter data have largely relied on keyword-matching or topic models for clustering relevant tweets. However, both methods suffer from the short-length of texts and unpredictable noise that naturally occurs in user-generated contexts. In response, we introduce a deep learning approach that uses hashtags as a form of supervision and learns tweet embeddings for extracting informative textual features. In this case study, we address the specific task of estimating state-level obesity from dietary-related textual features. Our approach yields an estimation that strongly correlates the textual features to government data and outperforms the keyword-matching baseline. The results also demonstrate the potential of discovering risk factors using the textual features. This method is general-purpose and can be applied to a wide range of Twitter-based public health studies.

Related