SPECTER: Document-level Representation Learning using Citation-informed Transformers
2020/04/15 by Arman Cohan, Cohan, Arman, Sergey Feldman +7 · 109 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
paper · pdf · doi:10.48550/arxiv.2004.07180
ACL 2020
arxiv created 2020/05/20 · arxiv updated 2020/05/21
Abstract
Representation learning is a critical ingredient for natural language processing systems. Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token- and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power. For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks. We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph. Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning. Additionally, to encourage further research on document-level models, we introduce SciDocs, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation. We show that SPECTER outperforms a variety of competitive baselines on the benchmark.
Cited by
- Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
- VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy
- Citation importance-aware document representation learning for large-scale science mapping
- Intelligent Scientific Literature Explorer using Machine Learning (ISLE)
- NoveltyRank: A Retrieval-Augmented Framework for Conceptual Novelty Estimation in AI Research
- Learned-Rule-Augmented Large Language Model Evaluators
- EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
- From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- FOS: A Large-Scale Temporal Graph Benchmark for Scientific Interdisciplinary Link Prediction
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
- BioMedJImpact: A Comprehensive Dataset and LLM Pipeline for AI Engagement and Scientific Impact Analysis of Biomedical Journals
- Practical Author Name Disambiguation under Metadata Constraints: A Contrastive Learning Approach for Astronomy Literature
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- CC30k: A Citation Contexts Dataset for Reproducibility-Oriented Sentiment Analysis
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- Contradictions in Context: Challenges for Retrieval-Augmented Generation in Healthcare
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- Discourse-Aware Scientific Paper Recommendation via QA-Style Summarization and Multi-Level Contrastive Learning
- Using language models to label clusters of scientific documents
- G2rammar: Bilingual Grammar Modeling for Enhanced Text-attributed Graph Learning
- Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
- Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?
- BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels
- DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
- Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
- SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
- ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative Retrieval
- Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
- Overview of the Plagiarism Detection Task at PAN 2025
- Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
- Contrastive Retrieval Heads Improve Attention-Based Re-Ranking
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- What Should I Cite? A RAG Benchmark for Academic Citation Prediction
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- An Artificial Intelligence Driven Semantic Similarity-Based Pipeline for Rapid Literature
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- MA-DPR: Manifold-aware Distance Metrics for Dense Passage Retrieval
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
- A Survey of Long-Document Retrieval in the PLM and LLM Era
- FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
- Parameter-Free Structural-Diversity Message Passing for Graph Neural Networks
- GeoGPT-RAG Technical Report
- CASPER: Concept-integrated Sparse Representation for Scientific Retrieval
- PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
- Uncovering drivers of climate research in policy with pretrained language models
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
- Uncertainty-driven Embedding Convolution
- SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search
- Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
- IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering
- Distillation versus Contrastive Learning: How to Train Your Rerankers
- From Ambiguity to Accuracy: The Transformative Effect of Coreference Resolution on Retrieval-Augmented Generation systems
- A Comparative Study of Specialized LLMs as Dense Retrievers
- Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences
- A Dynamical Cartography of the Epistemic Diffusion of Artificial Intelligence in Neuroscience
- SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
- Extracting Information About Publication Venues Using Citation-Informed Transformers
- Density, asymmetry and citation dynamics in scientific literature
- Literature-Grounded Novelty Assessment of Scientific Ideas
- ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research
- LGAI-EMBEDDING-Preview Technical Report
- A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
- Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem
- Aethorix v1.0: An Integrated Scientific AI Agent for Scalable Inorganic Materials Innovation and Industrial Implementation
- MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
- Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation
- QBD-RankedDataGen: Generating Custom Ranked Datasets for Improving Query-By-Document Search Using LLM-Reranking with Reduced Human Effort
- ERU-KG: Efficient Reference-aligned Unsupervised Keyphrase Generation
- MIR: Methodology Inspiration Retrieval for Scientific Research Problems
- Graph-Assisted Culturally Adaptable Idiomatic Translation for Indic Languages
- Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
- PD3: A Project Duplication Detection Framework via Adapted Multi-Agent Debate
- Internal and External Impacts of Natural Language Processing Papers
- Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
- Unify Graph Learning with Text: Unleashing LLM Potentials for Session Search
- Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- LDIR: Low-Dimensional Dense and Interpretable Text Embeddings with Relative Representations
- VizCV: AI-assisted visualization of researchers' publications tracks
- Benchmarking Retrieval-Augmented Generation for Chemistry
- MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
- SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- Towards a new paradigm of scientific discovery with socialized artificial intelligence
- SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG
- 100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
- Improving Scientific Document Retrieval with Academic Concept Index
- SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
- AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine
- CSPLADE: Learned Sparse Retrieval with Causal Language Models
- Goal Setting in Accounting Research: A Systematic Review and Reflections on Future Research Opportunities With AI‐Assisted Augmentation
- Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
- Causal Retrieval with Semantic Consideration
- Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
Related