2023/08/10 by Pedro Ruas, Diana Sousa, Ruas, Pedro +7
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Artificial intelligence #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #Computer science #Data science #Database #Engineering #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Information extraction #Named-entity recognition #Natural Language Processing Techniques #Natural language processing #Performance (cs.PF) #Pipeline (software) #Programming language #Relation (database) #Software #Software engineering #Systems engineering #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2308.05609
openalex publication_date 2023/08/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Biomedical Natural Language Processing (NLP) tends to become cumbersome for most researchers, frequently due to the amount and heterogeneity of text to be processed. To address this challenge, the industry is continuously developing highly efficient tools and creating more flexible engineering solutions. This work presents the integration between industry data engineering solutions for efficient data processing and academic systems developed for Named Entity Recognition (LasigeUnicage_NER) and Relation Extraction (BiOnt). Our design reflects an integration of those components with external knowledge in the form of additional training data from other datasets and biomedical ontologies. We used this pipeline in the 2022 LitCoin NLP Challenge, where our team LasigeUnicage was awarded the 7th Prize out of approximately 200 participating teams, reflecting a successful collaboration between the academia (LASIGE) and the industry (Unicage). The software supporting this work is available at \urlhttps://github.com/lasigeBioTM/Litcoin-LasigeUnicage.