2014/05/10 by Srayan Datta, Datta, Srayan
Computer Science · Physics and Astronomy · #Advanced Text Analysis Techniques #Complex Network Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Physical sciences #Information Retrieval (cs.IR) #Physics and Society (physics.soc-ph) #Social and Information Networks (cs.SI) #Web Data Mining and Analysis #cs.CL #cs.IR #cs.SI #physics.soc-ph
paper · pdf · doi:10.48550/arxiv.1405.2386
arxiv created 2014/05/10 · openalex publication_date 2014/05/10 · arxiv updated 2014/05/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In today's content-centric Internet, blogs are becoming increasingly popular and important from a data analysis perspective. According to Wikipedia, there were over 156 million public blogs on the Internet as of February 2011. Blogs are a reflection of our contemporary society. The contents of different blog posts are important from social, psychological, economical and political perspectives. Discovery of important topics in the blogosphere is an area which still needs much exploring. We try to come up with a procedure using probabilistic topic modeling and network centrality measures which identifies the central topics in a blog corpus.