vix.ing · top · new · best · stats

Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks

2026/01/01 by Atsuki Yamaguchi, Maggie Mi, Nikolaos Aletras · 1 citation
Arts and Humanities · Computer Science · Social Sciences · #EFL/ESL Teaching and Learning #Intelligent Tutoring Systems and Adaptive Learning #Writing and Handwriting Education #cs.CL

paper · pdf · doi:10.18653/v1/2026.acl-short.27

Accepted to ACL 2026 Main Conference

openalex publication_date 2026/01/01 · arxiv created 2026/04/15 · openalex created_date 2026/07/02 · openalex updated_date 2026/07/29 · arxiv updated 2026/07/30

Abstract

Language models (LMs) are pre-trained on raw text datasets to generate text sequences token-by-token. While this approach facilitates the learning of world knowledge and reasoning, it does not explicitly optimize for linguistic competence. To bridge this gap, we propose L2T, a pre-training framework integrating Language Learning Tasks alongside standard next-token prediction. Inspired by human language acquisition, L2T transforms raw text into structured input-output pairs to provide explicit linguistic stimulation. Pre-training LMs on a mixture of raw text and L2T data not only improves overall performance on linguistic competence benchmarks but accelerates its acquisition, while maintaining competitive performance on general reasoning tasks.

Citations

Cited by

Related