2024/03/22 by Thales Bertaglia, Bertaglia, Thales, Lily Heisig +5 · 1 citation
Business, Management and Accounting · Computer Science · #AI in Service Interactions #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FinTech, Crowdfunding, Digital Finance #Social and Information Networks (cs.SI) #Spam and Phishing Detection
paper · pdf · doi:10.48550/arxiv.2403.15214
openalex publication_date 2024/03/22 · openalex created_date 2024/03/26 · openalex updated_date 2026/07/28
Large Language Models (LLMs) raise concerns about lowering the cost of generating texts that could be used for unethical or illegal purposes, especially on social media. This paper investigates the promise of such models to help enforce legal requirements related to the disclosure of sponsored content online. We investigate the use of LLMs for generating synthetic Instagram captions with two objectives: The first objective (fidelity) is to produce realistic synthetic datasets. For this, we implement content-level and network-level metrics to assess whether synthetic captions are realistic. The second objective (utility) is to create synthetic data that is useful for sponsored content detection. For this, we evaluate the effectiveness of the generated synthetic data for training classifiers to identify undisclosed advertisements on Instagram. Our investigations show that the objectives of fidelity and utility may conflict and that prompt engineering is a useful but insufficient strategy. Additionally, we find that while individual synthetic posts may appear realistic, collectively they lack diversity, topic connectivity, and realistic user interaction patterns.