2025/06/09 by Shang Qu, Qu, Shang, Ning Ding +28
Biochemistry, Genetics and Molecular Biology · Decision Sciences · #Artificial Intelligence (cs.AI) #Bioinformatics and Genomic Networks #Biomedical Text Mining and Ontologies #FOS: Biological sciences #FOS: Computer and information sciences #Quantitative Methods (q-bio.QM) #Scientific Computing and Data Management
paper · pdf · doi:10.48550/arxiv.2506.07591
openalex publication_date 2025/06/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper introduces PROTEUS, a fully automated system that produces data-driven hypotheses from raw data files. We apply PROTEUS to clinical proteogenomics, a field where effective downstream data analysis and hypothesis proposal is crucial for producing novel discoveries. PROTEUS uses separate modules to simulate different stages of the scientific process, from open-ended data exploration to specific statistical analysis and hypothesis proposal. It formulates research directions, tools, and results in terms of relationships between biological entities, using unified graph structures to manage complex research processes. We applied PROTEUS to 10 clinical multiomics datasets from published research, arriving at 360 total hypotheses. Results were evaluated through external data validation and automatic open-ended scoring. Through exploratory and iterative research, the system can navigate high-throughput and heterogeneous multiomics data to arrive at hypotheses that balance reliability and novelty. In addition to accelerating multiomic analysis, PROTEUS represents a path towards tailoring general autonomous systems to specialized scientific domains to achieve open-ended hypothesis generation from data.