2025/10/01 by Cyril Pommier, Isabelle Alic, Llorenç Cabrera‐Bosquet +6 · 1 voice
Computer Science · Decision Sciences · Biochemistry, Genetics and Molecular Biology · #Research Data Management Practices #Scientific Computing and Data Management #Gene expression and cancer classification
paper · doi:10.1016/j.tplants.2025.09.001
openalex publication_date 2025/10/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Phenotypic datasets are increasingly rich and heterogeneous, with images, time courses, manual measurements, processed variables, and metadata. The management of such datasets navigates between partly incompatible objectives: (i) facilitate data analysis by extracting, organizing, and storing relevant variables; and (ii) allow reuse of raw, synthesized, and computed data (FAIR principles). For the first objective, 'dedicated datasets' can be extracted from raw information and tailored for the user's data analysis, but they result in a massive loss of information. We advocate that, for the second objective, 'sensu stricto phenomic datasets', upstream of dedicated datasets, should organize data without loss of information with data-science tools, in a 'theory-agnostic' way. They allow different users to build their own 'dedicated datasets' according to planned data analysis.