vix.ing · top · new · best · stats · spec

Topical Discovery of Web Content

2015/07/08 by Giancarlo Crocetti, Crocetti, Giancarlo
Computer Science · #Advanced Text Analysis Techniques #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Semantic Web and Ontologies #Web Data Mining and Analysis

paper · pdf · doi:10.48550/arxiv.1507.02002

openalex publication_date 2015/07/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This work describes the theory and the implementation of a new software tool, the "Web Topical Discovery System" (WTDS), which provides an approach to the automatic discovery and selection of new web pages relevant to specific analytical needs. We will see how it is possible to specify the research context with search keywords related to the area of interest and consider the important problem of removing extraneous data from a web page containing an article in order to reduce, to a minimum, false positives represented by a match on a keyword that is showing up on the latest news box of the same page. The removal of duplicates, the analysis of richness of information contained in the article and lexical diversity are all taken into consideration in order to provide the optimum set of recommendations to the end user or system.

Related