vix.ing · top · new · best · stats · spec

SAFE setup for generative molecular design

2024/10/26 by Yassir El Mesbahi, Mesbahi, Yassir El, Emmanuel Noutahi +1 · 1 voice · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Environmental Science · #Biomolecules (q-bio.BM) #Chemistry and Chemical Engineering #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG #q-bio.BM

paper · pdf · doi:10.48550/arxiv.2410.20232

openalex publication_date 2024/10/26 · arxiv published 2024/10/26 · arxiv updated 2024/10/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

SMILES-based molecular generative models have been pivotal in drug design but face challenges in fragment-constrained tasks. To address this, the Sequential Attachment-based Fragment Embedding (SAFE) representation was recently introduced as an alternative that streamlines those tasks. In this study, we investigate the optimal setups for training SAFE generative models, focusing on dataset size, data augmentation through randomization, model architecture, and bond disconnection algorithms. We found that larger, more diverse datasets improve performance, with the LLaMA architecture using Rotary Positional Embedding proving most robust. SAFE-based models also consistently outperform SMILES-based approaches in scaffold decoration and linker design, particularly with BRICS decomposition yielding the best results. These insights highlight key factors that significantly impact the efficacy of SAFE-based generative models.

Cited by

Discussions

Related