2025/08/06 by Aleksei Shestov, Shestov, Aleksei, Omar Zoloev +11
Computer Science · Decision Sciences · #Data Quality and Management #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Recommender Systems and Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2508.05688
openalex publication_date 2025/08/06 · openalex created_date 2025/10/15 · openalex updated_date 2026/07/28
This paper presents LLM4ES, a novel framework that exploits large pre-trained language models (LLMs) to derive user embeddings from event sequences. Event sequences are transformed into a textual representation, which is subsequently used to fine-tune an LLM through next-token prediction to generate high-quality embeddings. We introduce a text enrichment technique that enhances LLM adaptation to event sequence data, improving representation quality for low-variability domains. Experimental results demonstrate that LLM4ES achieves state-of-the-art performance in user classification tasks in financial and other domains, outperforming existing embedding methods. The resulting user embeddings can be incorporated into a wide range of applications, from user segmentation in finance to patient outcome prediction in healthcare.