2024/02/17 by Deepak Prasad, Prasad, Deepak, Mayur Pimpude +3 · 3 citations
Business, Management and Accounting · Computer Science · Engineering · #Business Process Modeling and Analysis #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Manufacturing Process and Optimization #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2402.11323
openalex publication_date 2024/02/17 · openalex created_date 2024/02/22 · openalex updated_date 2026/07/28
In this work a Large Language Model (LLM) based workflow is presented that utilizes OpenAI ChatGPT model GPT-3.5-turbo-1106 and Google Gemini Pro model to create summary of text, data and images from research articles. It is demonstrated that by using a series of processing, the key information can be arranged in tabular form and knowledge graphs to capture underlying concepts. Our method offers efficiency and comprehension, enabling researchers to extract insights more effectively. Evaluation based on a diverse Scientific Paper Collection demonstrates our approach in facilitating discovery of knowledge. This work contributes to accelerated material design by smart literature review. The method has been tested based on various qualitative and quantitative measures of gathered information. The ChatGPT model achieved an F1 score of 0.40 for an exact match (ROUGE-1, ROUGE-2) but an impressive 0.479 for a relaxed match (ROUGE-L, ROUGE-Lsum) structural data format in performance evaluation. The Google Gemini Pro outperforms ChatGPT with an F1 score of 0.50 for an exact match and 0.63 for a relaxed match. This method facilitates high-throughput development of a database relevant to materials informatics. For demonstration, an example of data extraction and knowledge graph formation based on a manuscript about a titanium alloy is discussed.