Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications
2025/05/20 by Fadel M. Megahed, Megahed, Fadel M., Ying‐Ju Chen +12 · 1 voice · 1 citation
Computer Science · Mathematics · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Text and Document Classification Technologies #cs.CL #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2505.14918
openalex publication_date 2025/05/20 · arxiv published 2025/05/20 · openalex created_date 2025/10/19 · arxiv updated 2025/12/19 · openalex updated_date 2026/07/28
Abstract
This study introduces a framework for evaluating consistency in large language model (LLM) binary text classification, addressing the lack of established reliability assessment methods. Adapting psychometric principles, we determine sample size requirements, develop metrics for invalid responses, and evaluate intra- and inter-rater reliability. Our case study examines financial news sentiment classification across 14 LLMs (including claude-3-7-sonnet, gpt-4o, deepseek-r1, gemma3, llama3.2, phi4, and command-r-plus), with five replicates per model on 1,350 articles. Models demonstrated high intra-rater consistency, achieving perfect agreement on 90-98% of examples, with minimal differences between expensive and economical models from the same families. When validated against StockNewsAPI labels, models achieved strong performance (accuracy 0.76-0.88), with smaller models like gemma3:1B, llama3.2:3B, and claude-3-5-haiku outperforming larger counterparts. All models performed at chance when predicting actual market movements, indicating task constraints rather than model limitations. Our framework provides systematic guidance for LLM selection, sample size planning, and reliability assessment, enabling organizations to optimize resources for classification tasks.
Citations
- Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
- Scaling Public Health Text Annotation: Zero-Shot Learning vs. Crowdsourcing for Improved Efficiency and Labeling Accuracy
- Adapting OpenAI's CLIP Model for Few-Shot Image Inspection in Manufacturing Quality Control: An Expository Case Study with Multiple Application Examples
- Large Language Models For Text Classification: Case Study And Comprehensive Review
- Revisiting time-varying dynamics in stock market forecasting: A multi-source sentiment analysis approach with large language model
- A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges
- Large Language Models are Inconsistent and Biased Evaluators
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Large Language Models: A Survey
- Best Practices for Text Annotation with Large Language Models
- Prompt Design and Engineering: Introduction and Advanced Methods
- Entity Matching using Large Language Models
- Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning
- Automated Annotation with Generative AI Requires Validation
- Text Classification via Large Language Models
- Can Large Language Models Be an Alternative to Human Evaluations?
- Testing the Reliability of ChatGPT for Text Annotation and Classification: A Cautionary Remark
- ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning
- ChatGPT outperforms crowd workers for text-annotation tasks
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations
- Want To Reduce Labeling Cost? GPT-3 Can Help
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Breaking Community Boundary: Comparing Academic and Social Communication Preferences regarding Global Pandemics
- Language Models are Few-Shot Learners
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- Computing inter‐rater reliability and its variance in the presence of high agreement
- Coefficient Kappa: Some Uses, Misuses, and Alternatives
- The Measurement of Observer Agreement for Categorical Data
- A Coefficient of Agreement for Nominal Scales
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Discussions
Related