2024/08/15 by David F. Farr, Farr, David, Nico Manzonelli +5
Social Sciences · #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Social and Information Networks (cs.SI)
paper · pdf · doi:10.48550/arxiv.2408.08217
openalex publication_date 2024/08/15 · openalex created_date 2025/01/03 · openalex updated_date 2026/07/28
Large language models (LLMs) have enhanced our ability to rapidly analyze and classify unstructured natural language data. However, concerns regarding cost, network limitations, and security constraints have posed challenges for their integration into work processes. In this study, we adopt a systems design approach to employing LLMs as imperfect data annotators for downstream supervised learning tasks, introducing novel system intervention measures aimed at improving classification performance. Our methodology outperforms LLM-generated labels in seven of eight tests, demonstrating an effective strategy for incorporating LLMs into the design and deployment of specialized, supervised learning models present in many industry use cases.