2016/06/28 by Steven Kearnes, Kearnes, Steven, Brian Goldman +3 · 2 citations
Computer Science · Engineering · Materials Science · #Computational Drug Discovery Methods #FOS: Computer and information sciences #Innovative Microfluidic and Catalytic Techniques Innovation #Machine Learning (stat.ML) #Machine Learning in Materials Science
paper · pdf · doi:10.48550/arxiv.1606.08793
openalex publication_date 2016/06/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Deep learning methods such as multitask neural networks have recently been applied to ligand-based virtual screening and other drug discovery applications. Using a set of industrial ADMET datasets, we compare neural networks to standard baseline models and analyze multitask learning effects with both random cross-validation and a more relevant temporal validation scheme. We confirm that multitask learning can provide modest benefits over single-task models and show that smaller datasets tend to benefit more than larger datasets from multitask learning. Additionally, we find that adding massive amounts of side information is not guaranteed to improve performance relative to simpler multitask learning. Our results emphasize that multitask effects are highly dataset-dependent, suggesting the use of dataset-specific models to maximize overall performance.