vix.ing · top · new · best · stats · spec

Relation Extraction in underexplored biomedical domains: A diversity-optimised sampling and synthetic data generation approach

2023/11/10 by Delmas, Maxime, Magdalena Wysocka, Wysocka, Magdalena +2
Biochemistry, Genetics and Molecular Biology · Medicine · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Microbial Natural Products and Biosynthesis

paper · pdf · doi:10.48550/arxiv.2311.06364

openalex publication_date 2023/11/10 · openalex created_date 2023/11/15 · openalex updated_date 2026/07/28

Abstract

A new synthetic dataset (training/validation) for end-to-end Relation Extraction of relationships between Organisms and Natural-Products. The new dataset was generated using Mixtral-8x7B-Instruct-v0.1. Like the model, the produced synthetic data are also submitted to the License of the model used for generation (apache-2.0). The new dataset was created based on the top-1000 (per biological kingdom) LOTUS literature references extracted with the GME-sampler. The dataset contains 8,913 items in the training set and 344 items in the validation set. The dataset was generated using the same protocol as described in the article.

Related