vix.ing · top · new · best · stats · spec

Prompt Refinement or Fine-tuning? Best Practices for using LLMs in Computational Social Science Tasks

2024/08/02 by Anders Giovanni Møller, Luca Maria Aiello, Møller, Anders Giovanni +1 · 1 voice · 1 citation
Social Sciences · #Computational and Text Analysis Methods #cs.CL #cs.CY #physics.soc-ph

paper · pdf · doi:10.48550/arxiv.2408.01346

openalex publication_date 2024/08/02 · openalex created_date 2025/01/13 · openalex updated_date 2026/07/28

Abstract

Large Language Models are expressive tools that enable complex tasks of text understanding within Computational Social Science. Their versatility, while beneficial, poses a barrier for establishing standardized best practices within the field. To bring clarity on the values of different strategies, we present an overview of the performance of modern LLM-based classification methods on a benchmark of 23 social knowledge tasks. Our results point to three best practices: select models with larger vocabulary and pre-training corpora; avoid simple zero-shot in favor of AI-enhanced prompting; fine-tune on task-specific data, and consider more complex forms instruction-tuning on multiple datasets only when only training data is more abundant.

Cited by

Discussions

Related