2015/07/08 by Giancarlo Crocetti, Crocetti, Giancarlo
Computer Science · #Advanced Text Analysis Techniques #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Semantic Web and Ontologies #Web Data Mining and Analysis
paper · pdf · doi:10.48550/arxiv.1507.02002
openalex publication_date 2015/07/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This work describes the theory and the implementation of a new software tool, the "Web Topical Discovery System" (WTDS), which provides an approach to the automatic discovery and selection of new web pages relevant to specific analytical needs. We will see how it is possible to specify the research context with search keywords related to the area of interest and consider the important problem of removing extraneous data from a web page containing an article in order to reduce, to a minimum, false positives represented by a match on a keyword that is showing up on the latest news box of the same page. The removal of duplicates, the analysis of richness of information contained in the article and lexical diversity are all taken into consideration in order to provide the optimum set of recommendations to the end user or system.