vix.ing · top · new · best · stats · spec

Automating biomedical data science through tree-based pipeline\n optimization

2016/01/28 by Randal S. Olson, Ryan J. Urbanowicz, Olson, Randal S. +9 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Multi-Objective Optimization Algorithms #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Machine Learning in Bioinformatics #Neural and Evolutionary Computing (cs.NE)

paper · pdf · doi:10.48550/arxiv.1601.07925

openalex publication_date 2016/01/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Over the past decade, data science and machine learning has grown from a\nmysterious art form to a staple tool across a variety of fields in academia,\nbusiness, and government. In this paper, we introduce the concept of tree-based\npipeline optimization for automating one of the most tedious parts of machine\nlearning---pipeline design. We implement a Tree-based Pipeline Optimization\nTool (TPOT) and demonstrate its effectiveness on a series of simulated and\nreal-world genetic data sets. In particular, we show that TPOT can build\nmachine learning pipelines that achieve competitive classification accuracy and\ndiscover novel pipeline operators---such as synthetic feature\nconstructors---that significantly improve classification accuracy on these data\nsets. We also highlight the current challenges to pipeline optimization, such\nas the tendency to produce pipelines that overfit the data, and suggest future\nresearch paths to overcome these challenges. As such, this work represents an\nearly step toward fully automating machine learning pipeline design.\n

Citations

Cited by

Related