2010/12/03 by Pekka Malo, Pyry Siitari, Malo, Pekka +7
Computer Science · Social Sciences · #Artificial intelligence #Computer science #Content (measure theory) #Domain (mathematical analysis) #Domain knowledge #FOS: Computer and information sciences #H.3.1 #H.3.3 #Information Retrieval (cs.IR) #Information retrieval #Natural language processing #Ontology #Semantic similarity #Support vector machine #Task (project management) #Text and Document Classification Technologies #Topic Modeling #Wikis in Education and Collaboration #cs.IR
paper · pdf · doi:10.48550/arxiv.1012.0854
9 pages, Third International Workshop on Semantic Aspects in Data Mining (SADM'10) in conjunction with the 2010 IEEE International Conference on Data Mining
arxiv created 2010/12/03 · openalex publication_date 2010/12/03 · arxiv updated 2015/03/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as medical free texts or legal documents. However, the problem is that sources of domain knowledge are time-consuming to build and equally costly to maintain. As a potential remedy, recent studies on Wikipedia suggest that this large body of socially constructed knowledge can be effectively harnessed to provide not only facts but also accurate information about semantic concept-similarities. This paper describes a framework for document filtering, where Wikipedia's concept-relatedness information is combined with a domain ontology to produce semantic content classifiers. The approach is evaluated using Reuters RCV1 corpus and TREC-11 filtering task definitions. In a comparative study, the approach shows robust performance and appears to outperform content classifiers based on Support Vector Machines (SVM) and C4.5 algorithm.