2024/11/07 by Sonal Prabhune, Prabhune, Sonal, Donald J. Berndt +1 · 1 voice · 2 citations
Computer Science · #68T01 #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.0 #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.IR
paper · pdf · doi:10.48550/arxiv.2411.11895
openalex publication_date 2024/11/07 · arxiv published 2024/11/07 · arxiv updated 2024/11/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Knowing that the generative capabilities of large language models (LLM) are sometimes hampered by tendencies to hallucinate or create non-factual responses, researchers have increasingly focused on methods to ground generated outputs in factual data. Retrieval Augmented Generation (RAG) has emerged as a key approach for integrating knowledge from data sources outside of the LLM's training set, including proprietary and up-to-date information. While many research papers explore various RAG strategies, their true efficacy is tested in real-world applications with actual data. The journey from conceiving an idea to actualizing it in the real world is a lengthy process. We present insights from the development and field-testing of a pilot project that integrates LLMs with RAG for information retrieval. Additionally, we examine the impacts on the information value chain, encompassing people, processes, and technology. Our aim is to identify the opportunities and challenges of implementing this emerging technology, particularly within the context of behavioral research in the information systems (IS) field. The contributions of this work include the development of best practices and recommendations for adopting this promising technology while ensuring compliance with industry regulations through a proposed AI governance model.