2017/11/13 by Philippe Schwaller, Schwaller, Philippe, Théophile Gaudin +7 · 4 citations
Computer Science · Materials Science · #Advanced Text Analysis Techniques #Computational Drug Discovery Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Materials Science #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1711.04810
openalex publication_date 2017/11/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
There is an intuitive analogy of an organic chemist's understanding of a\ncompound and a language speaker's understanding of a word. Consequently, it is\npossible to introduce the basic concepts and analyze potential impacts of\nlinguistic analysis to the world of organic chemistry. In this work, we cast\nthe reaction prediction task as a translation problem by introducing a\ntemplate-free sequence-to-sequence model, trained end-to-end and fully\ndata-driven. We propose a novel way of tokenization, which is arbitrarily\nextensible with reaction information. With this approach, we demonstrate\nresults superior to the state-of-the-art solution by a significant margin on\nthe top-1 accuracy. Specifically, our approach achieves an accuracy of 80.1%\nwithout relying on auxiliary knowledge such as reaction templates. Also, 66.4%\naccuracy is reached on a larger and noisier dataset.\n