vix.ing · top · new · best · stats · spec

Multilingual Medical Documents Classification Based on MesH Domain\n Ontology

2012/06/21 by Zakaria Elberrichi, Elberrichi, Zakaria, Malika Taibi +3
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Text Analysis Techniques #Biomedical Text Mining and Ontologies #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Text and Document Classification Technologies

paper · pdf · doi:10.48550/arxiv.1206.4883

openalex publication_date 2012/06/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This article deals with the semantic Web and ontologies. It addresses the\nissue of the classification of multilingual Web documents, based on domain\nontology. The objective is being able, using a model, to classify documents in\ndifferent languages. We will try to solve this problematic using two different\napproaches. The two approaches will have two elementary stages: the creation of\nthe model using machine learning algorithms on a labeled corpus, then the\nclassification of documents after detecting their languages and mapping their\nterms into the concepts of the language of reference (English). But each one\nwill deal with the multilingualism with a different approach. One supposes the\nontology is monolingual, whereas the other considers it multilingual. To show\nthe feasibility and the importance of our work, we implemented it on a domain\nthat attracts nowadays a lot of attention from the data mining community: the\nbiomedical domain. The selected documents are from the biomedical benchmark\ncorpus Ohsumed, and the associated ontology is the thesaurus MeSH (Medical\nSubject Headings). The main idea in our work is a new document representation,\nthe masterpiece of all good classification, based on concept. The experimental\nresults show that the recommended ideas are promising.\n

Related