2025/05/11 by Yosuke Kikuchi, Kikuchi, Yosuke, Yasuhiro Yoshikai +11
Biochemistry, Genetics and Molecular Biology · Computer Science · Materials Science · #92E10 #Biomedical Text Mining and Ontologies #Computational Drug Discovery Methods #FOS: Biological sciences #I.2.6 #J.3 #Machine Learning in Materials Science #Quantitative Methods (q-bio.QM)
paper · pdf · doi:10.48550/arxiv.2505.07139
openalex publication_date 2025/05/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Chemical language models (CLMs) are increasingly used for molecular design and property prediction. Because these models learn from textual encodings of molecules, differences in how such encodings are generated may affect their behavior. In cheminformatics, the term canonical SMILES implies a single standardized notation, yet different toolkits define distinct canonicalization rules, yielding multiple canonical strings for the same molecule. To examine how this variability arises and why it matters, we surveyed 264 CLM papers in PubMed and found that about half did not specify their canonicalization procedure, limiting transparency and reproducibility. Using a molecular translation framework, we show that when multiple valid notations are mixed or left undocumented, inconsistent notations distort latent representations and, in some benchmarks, can spuriously inflate predictive accuracy, a phenomenon we term notation-level confounding. These findings demonstrate how subtle differences in SMILES generation can mislead CLMs and highlight the importance of explicitly reporting preprocessing tools and settings.