2025/04/12 by Lennart Finke, Chandan Sreedhara, Finke, Lennart +18 · 1 voice · 1 citation
Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2504.09184
openalex publication_date 2025/04/12 · openalex created_date 2025/10/14 · openalex updated_date 2026/07/28
We present SimpleStories, a large synthetic story dataset in simple language, consisting of 2 million samples each in English and Japanese. Through parameterizing prompts at multiple levels of abstraction, we achieve control over story characteristics at scale, inducing syntactic and semantic diversity. Ablations on a newly trained model suite show improved sample efficiency and model interpretability compared to the TinyStories dataset. We open-source all constituent parts of model creation, hoping to enable novel ways to study the end-to-end training process. As a byproduct, we move the frontier regarding the fewest-parameter language model that outputs grammatical natural language.