vix.ing · top · new · best · stats · spec

LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

2025/04/20 by Xia, Yunhui, Shen, Wei, Wang, Yan +5 · 18 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering (cs.SE)

paper · doi:10.48550/arxiv.2504.14655

Abstract

We introduce LeetCodeDataset, a high-quality benchmark for evaluating and training code-generation models, addressing two key challenges in LLM research: the lack of reasoning-focused coding benchmarks and self-contained training testbeds. By curating LeetCode Python problems with rich metadata, broad coverage, 100+ test cases per problem, and temporal splits (pre/post July 2024), our dataset enables contamination-free evaluation and efficient supervised fine-tuning (SFT). Experiments show reasoning models significantly outperform non-reasoning counterparts, while SFT with only 2.6K model-generated solutions achieves performance comparable to 110K-sample counterparts. The dataset and evaluation framework are available on Hugging Face and Github.

Cited by

Related