2025/04/09 by Jiacheng Liu, Liu, Jiacheng, Taylor Blanton +61 · 3 voices · 16 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Data modeling #Language identification #Language model #Natural Language Processing Techniques #Natural language #Topic Modeling #Tracing #Training (meteorology) #Training set #Universal Networking Language
paper · pdf · doi:10.48550/arxiv.2504.07096
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/04/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We present OLMoTrace, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace finds and shows verbatim matches between segments of language model output and documents in the training text corpora. Powered by an extended version of infini-gram (Liu et al., 2024), our system returns tracing results within a few seconds. OLMoTrace can help users understand the behavior of language models through the lens of their training data. We showcase how it can be used to explore fact checking, hallucination, and the creativity of language models. OLMoTrace is publicly available and fully open-source.