vix.ing · top · new · best · stats · spec

Chasing Random: Instruction Selection Strategies Fail to Generalize

2024/10/19 by Harshita Diddee, Daphne Ippolito, Diddee, Harshita +1 · 2 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Educational Assessment and Pedagogy #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Online and Blended Learning

paper · pdf · doi:10.48550/arxiv.2410.15225

openalex publication_date 2024/10/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions. This has accelerated the development of methods that filter a large, noisy instruction-tuning datasets down to high-quality subset which works just as well. However, typically, the performance of these methods is not demonstrated across a uniform experimental setup and thus their generalization capabilities are not well established. In this work, we analyze popular selection strategies across different source datasets, selection budgets and evaluation benchmarks: Our results indicate that selection strategies generalize poorly, often failing to consistently outperform even random baselines. We also analyze the cost-performance trade-offs of using data selection. Our findings reveal that data selection can often exceed the cost of fine-tuning on the full dataset, yielding only marginal and sometimes no gains compared to tuning on the full dataset or a random subset.

Cited by

Related