2025/06/11 by Simin Ma, Zeng, Qingyun, Ma, Simin +6
Computer Science · #Advanced Database Systems and Queries #Natural Language Processing Techniques #Logic, programming, and type systems
paper · pdf · doi:10.48550/arxiv.2506.09359
The rise of Large Language Models (LLMs) has significantly advanced Text-to-SQL (NL2SQL) systems, yet evaluating the semantic equivalence of generated SQL remains a challenge, especially given ambiguous user queries and multiple valid SQL interpretations. This paper explores using LLMs to assess both semantic and a more practical "weak" semantic equivalence. We analyze common patterns of SQL equivalence and inequivalence, discuss challenges in LLM-based evaluation.