vix.ing · top · new · best · stats · spec

Do humans and large language models agree on the quality of synthesis plans?

2026/04/08 by Varvara Voinarovska, Roćıo Mercado, Mikhail Kabeshov +1 · 1 voice
Materials Science · Social Sciences · Environmental Science · #Machine Learning in Materials Science #Language and cultural evolution #Chemistry and Chemical Engineering

paper · doi:10.26434/chemrxiv.15001730/v1

Abstract

Large language models (LLMs) have seen a widespread adoption in all spheres of science including chemistry and cheminformatics. Nevertheless, our knowledge of how they operate is limited, giving rise to exploration of their capabilities in different areas of science and different operation modes. Here, we investigated whether LLMs could mimic human experts on the challenging task of assessing retrosynthetic path feasibility in routes generated by a popular computer-aided synthesis planning tool (AiZynthFinder). We evaluated the agreement between LLMs and expert chemists on holistic evaluations of the proposed routes as well as the individual chemical reactions in them. We used four frontier LLMs (Claude Sonnet 4.5, GPT-4.1, GPT-o3, and Gemini 2.5 Pro) and employed 17 expert chemists to grade 50 retrosynthetic paths. We found out that when instructed with clear possible categories for the evaluation of a given reaction in a retrosynthetic tree, human experts tend to converge in their opinions. Another important finding was that Gemini 2.5 Pro performs remarkably well out of all the LLMs explored, GPT-o3 tends to be more pessimistic, and Claude and GPT-4.1 tend to be overly optimistic when compared to the majority vote of human experts.

Discussions

Related