2025/06/04 by Bowen Zheng, Grace X. Gu · 1 voice
Computer Science · #Topic Modeling #Advanced Text Analysis Techniques #Natural Language Processing Techniques
paper · pdf · doi:10.1002/aidi.202500036
Machine learning interatomic potential (MLIP) is an emerging technique that has helped achieve molecular dynamics simulations with unprecedented balance between efficiency and accuracy. Recently, the body of MLIP literature has been growing rapidly, which propels the need to automatically process relevant information for researchers to understand and utilize. Named entity recognition (NER), a natural language processing technique that identifies and categorizes information from texts, may help summarize key approaches and findings of relevant papers. In this work, we develop an NER model for MLIP literature by fine‐tuning a pre‐trained language model. To streamline text annotation, we build a user‐friendly web application for annotation and proofreading, which is seamlessly integrated into the training procedure. Our model can identify technical entities with an F1 score of 0.8 for new MLIP paper abstracts using only 60 training paper abstracts and up to 0.75 for scientific texts on different topics. Notably, some “errors” in predictions are actually reasonable decisions, showcasing the model's ability beyond what the performance metrics indicate. This work demonstrates the linguistic capabilities of the NER approach in processing textual information of a specific scientific domain and has the potential to accelerate materials research using language models and contribute to a user‐centric workflow.