vix.ing · top · new · best · stats · spec

Improving Vietnamese Legal Document Retrieval using Synthetic Data

2024/12/01 by Son Pham Tien, Hieu Nguyen Doan, Tien, Son Pham +4 · 1 citation
Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #FOS: Computer and information sciences #Information Retrieval (cs.IR)

paper · pdf · doi:10.48550/arxiv.2412.00657

openalex publication_date 2024/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In the field of legal information retrieval, effective embedding-based models are essential for accurate question-answering systems. However, the scarcity of large annotated datasets poses a significant challenge, particularly for Vietnamese legal texts. To address this issue, we propose a novel approach that leverages large language models to generate high-quality, diverse synthetic queries for Vietnamese legal passages. This synthetic data is then used to pre-train retrieval models, specifically bi-encoder and ColBERT, which are further fine-tuned using contrastive loss with mined hard negatives. Our experiments demonstrate that these enhancements lead to strong improvement in retrieval accuracy, validating the effectiveness of synthetic data and pre-training techniques in overcoming the limitations posed by the lack of large labeled datasets in the Vietnamese legal domain.

Cited by

Related