2016/06/24 by Richard J. Oentaryo, Oentaryo, Richard J., Ee-Peng Lim +8 · 4 citations
Computer Science · #Advanced Graph Neural Networks #Artificial intelligence #Computer science #Data science #FOS: Computer and information sciences #Human–computer interaction #Machine Learning (cs.LG) #Profiling (computer programming) #Social and Information Networks (cs.SI) #Social media #Spam and Phishing Detection #Text and Document Classification Technologies #World Wide Web #cs.LG #cs.SI
paper · pdf · doi:10.48550/arxiv.1606.07707
published in arXiv (Cornell University) (Cornell University)
arxiv created 2016/06/24 · openalex publication_date 2016/06/24 · arxiv updated 2016/06/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The abundance of user-generated data in social media has incentivized the development of methods to infer the latent attributes of users, which are crucially useful for personalization, advertising and recommendation. However, the current user profiling approaches have limited success, due to the lack of a principled way to integrate different types of social relationships of a user, and the reliance on scarcely-available labeled data in building a prediction model. In this paper, we present a novel solution termed Collective Semi-Supervised Learning (CSL), which provides a principled means to integrate different types of social relationship and unlabeled data under a unified computational framework. The joint learning from multiple relationships and unlabeled data yields a computationally sound and accurate approach to model user attributes in social media. Extensive experiments using Twitter data have demonstrated the efficacy of our CSL approach in inferring user attributes such as account type and marital status. We also show how CSL can be used to determine important user features, and to make inference on a larger user population.