vix.ing · top · new · best · stats · spec

What Makes You CLIC: Detection of Croatian Clickbait Headlines

2025/07/18 by Marija Anđelić, Anđelić, Marija, Dominik Šipek +5
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Radio, Podcasts, and Digital Media #Web Data Mining and Analysis

paper · pdf · doi:10.48550/arxiv.2507.14314

openalex publication_date 2025/07/18 · openalex created_date 2025/10/16 · openalex updated_date 2026/07/28

Abstract

Online news outlets operate predominantly on an advertising-based revenue model, compelling journalists to create headlines that are often scandalous, intriguing, and provocative -- commonly referred to as clickbait. Automatic detection of clickbait headlines is essential for preserving information quality and reader trust in digital media and requires both contextual understanding and world knowledge. For this task, particularly in less-resourced languages, it remains unclear whether fine-tuned methods or in-context learning (ICL) yield better results. In this paper, we compile CLIC, a novel dataset for clickbait detection of Croatian news headlines spanning a 20-year period and encompassing mainstream and fringe outlets. We fine-tune the BERTić model on this task and compare its performance to LLM-based ICL methods with prompts both in Croatian and English. Finally, we analyze the linguistic properties of clickbait. We find that nearly half of the analyzed headlines contain clickbait, and that finetuned models deliver better results than general LLMs.

Citations

Related