vix.ing · top · new · best · stats

On the Use of ArXiv as a Dataset

2019/04/30 by Colin B. Clement, Clement, Colin B., Matthew Bierbaum +6 · 23 citations
Computer Science · Decision Sciences · Physics and Astronomy · #Advanced Graph Neural Networks #FOS: Computer and information sciences #FOS: Physical sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Physics and Society (physics.soc-ph) #Scientific Computing and Data Management #Social and Information Networks (cs.SI) #Topic Modeling #cs.IR #cs.LG #cs.SI #physics.soc-ph

paper · pdf · doi:10.48550/arxiv.1905.00075

7 pages, 3 tables, 2 figures, ICLR 2019 workshop RLGM submission

arxiv created 2019/04/30 · openalex publication_date 2019/04/30 · arxiv updated 2019/05/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The arXiv has collected 1.5 million pre-print articles over 28 years, hosting literature from scientific fields including Physics, Mathematics, and Computer Science. Each pre-print features text, figures, authors, citations, categories, and other metadata. These rich, multi-modal features, combined with the natural graph structure---created by citation, affiliation, and co-authorship---makes the arXiv an exciting candidate for benchmarking next-generation models. Here we take the first necessary steps toward this goal, by providing a pipeline which standardizes and simplifies access to the arXiv's publicly available data. We use this pipeline to extract and analyze a 6.7 million edge citation graph, with an 11 billion word corpus of full-text research articles. We present some baseline classification results, and motivate application of more exciting generative graph models.

Citations

Cited by

Related