vix.ing · top · new · best · stats · spec

Benchmarking ESG Risk Classification in 10-K Filings

2026/01/22 by Ace Vo, Rosemary Kim, Miloslava Plachkinova +3
Engineering · Business, Management and Accounting · Computer Science · #Marine and Offshore Engineering Studies #Law, logistics, and international trade #Diverse Research and Applications

paper · doi:10.1080/08874417.2026.2619941

Abstract

This study systematically evaluates machine learning (ML) and large language models (LLMs) for classifying Environmental, Social, and Governance (ESG) risks in U.S. 10-K filings. Using the FinBERT-ESG-9-Categories model, fine-tuned on ESG disclosures, we classify over 122,000 paragraphs from S&P 100 firms’ Item 1A Risk Factor sections (2012–2022). A human-annotated dataset of 789 paragraphs benchmarks FinBERT against five LLMs: ChatGPT 4o, Claude 3.5 Sonnet, Grok 3, Gemini 2.5 Pro, and DeepSeek 3.1. FinBERT achieves the highest accuracy (83%) and macro F1-score (0.83), outperforming LLMs (59–69%). While LLMs show strong inter-model agreement, alignment with human labels is moderate, revealing limits of model consensus. LLMs also struggle with nuanced ESG categories like Community Relations and Business Ethics. Findings indicate domain-specific models outperform general LLMs for structured ESG analysis. A hybrid framework combining ML, LLM, and human review is recommended to improve accuracy and interpretability.

Citations

Related