SQuAD: 100,000+ Questions for Machine Comprehension of Text
2016/06/16 by Rajpurkar, Pranav, Zhang, Jian, Lopyrev, Konstantin +1 · 311 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.1606.05250
Abstract
We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text from the corresponding reading passage. We analyze the dataset to understand the types of reasoning required to answer the questions, leaning heavily on dependency and constituency trees. We build a strong logistic regression model, which achieves an F1 score of 51.0%, a significant improvement over a simple baseline (20%). However, human performance (86.8%) is much higher, indicating that the dataset presents a good challenge problem for future research. The dataset is freely available at https://stanford-qa.com
Cited by
- Rethinking the Capability of Fine-Tuned Language Models for Automated Vulnerability Repair
- Towards High-Level Semantic Intelligence
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Cross-Attention Calibrated Deduplication for Retrieval-Augmented Generation System
- Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
- cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs
- Model Merging via Multi-Teacher Knowledge Distillation
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- MEPIC: Memory Efficient Position Independent Caching for LLM Serving
- Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
- Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- ModelTables: A Corpus of Tables about Models
- Bolmo: Byteifying the Next Generation of Language Models
- Multiscale Aggregated Hierarchical Attention (MAHA): A Game Theoretic and Optimization Driven Approach to Efficient Contextual Modeling in Large Language Models
- Large-Language Memorization During the Classification of United States Supreme Court Cases
- FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
- Detecting Prompt Injection Attacks Against Application Using Classifiers
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
- CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- Enhancing Next-Generation Language Models with Knowledge Graphs: Extending Claude, Mistral IA, and GPT-4 via KG-BERT
- AgriRegion: Region-Aware Retrieval for High-Fidelity Agricultural Advice
- When unlearning is free: leveraging low influence points to reduce computational costs
- Towards a Science of Scaling Agent Systems
- Progress Ratio Embeddings: An Impatience Signal for Robust Length Control in Neural Text Generation
- ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
- Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale
- Learning Steerable Clarification Policies with Collaborative Self-play
- Lumos: Let there be Language Model System Certification
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
- Tourism Question Answer System in Indian Language using Domain-Adapted Foundation Models
- SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
- LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
- BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- Interactive AI NPCs Powered by LLMs: Technical Report for the CPDC Challenge 2025
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- A Systematic Study of Compression Ordering for Large Language Models
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
- Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
- PersonalizedRouter: Personalized LLM Routing via Graph-based User Preference Modeling
- Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
- HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- SweeperBot: Making 3D Browsing Accessible through View Analysis and Visual Question Answering
- Exploring question answering: metric analysis and evaluation framework for enhanced interpretability
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
- Local Hybrid Retrieval-Augmented Document QA
- Chain of Summaries: Summarization Through Iterative Questioning
- End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
- | \circlearrowright \boxedBUS |: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
- Testing Question Answering Software with Context-Driven Question Generation
- Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
- Differentially Private In-Context Learning with Nearest Neighbor Search
- ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
- Cache Mechanism for Agent RAG Systems
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
- KV Cache Transform Coding for Compact Storage in LLM Inference
- Toward Sustainability-Aware LLM Inference on Edge Clusters
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- Elastic Architecture Search for Efficient Language Models
- Integrating Ontologies with Large Language Models for Enhanced Control Systems in Chemical Engineering
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
- Metis: Memory Foundation Model
- Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling
- Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
- A Survey on Unlearning in Large Language Models
- Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- ChessQA: Evaluating Large Language Models for Chess Understanding
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
- SALSA: Single-pass Autoregressive LLM Structured Classification
- GigaEmbeddings: Efficient Russian Language Embedding Model
- Model Merging with Functional Dual Anchors
- Estonian Native Large Language Model Benchmark
- Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Simple Context Compression: Mean-Pooling and Multi-Ratio Training
- ARC-Encoder: learning compressed text representations for large language models
- LM-mixup: Text Data Augmentation via Language Model based Mixup
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
- LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
- IMB: An Italian Medical Benchmark for Question Answering
- Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
- Assessing Monotone Dependence: Area Under the Curve Meets Rank Correlation
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- Rethinking On-policy Optimization for Query Augmentation
- Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
- Annotation-Efficient Universal Honesty Alignment
- Vocab Diet: Reshaping the Vocabulary of LLMs with Vector Arithmetic
- Structured yet Bounded Temporal Understanding in Large Language Models
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- End-to-end Listen, Look, Speak and Act
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
- TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
- MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
- Understanding the Ability of LLMs to Handle Character-Level Perturbation
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
- BitNet Distillation
- Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
- K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
- Selective Adversarial Attacks on LLM Benchmarks
- Probing Latent Knowledge Conflict for Faithful Retrieval-Augmented Generation
- Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
- The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
- Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
- Doc-to-Atom: Learning to Compile and Compose Memory Atoms
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
- Revisiting Hallucination Detection with Effective Rank-based Uncertainty
- A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning
- Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence
- Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
- When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
- PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
- YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video
- SEER: The Span-based Emotion Evidence Retrieval Benchmark
- Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
- AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Learning to Route: A Rule-Driven Agent Framework for Hybrid-Source Retrieval-Augmented Generation
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- Accelerating LLM Inference with Precomputed Query Storage
- Mem-α: Learning Memory Construction via Reinforcement Learning
- ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
- ReTAG: Retrieval-Enhanced, Topic-Augmented Graph-Based Global Sensemaking
- Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
- Vision Function Layer in Multimodal LLMs
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- LLM DNA: Tracing Model Evolution via Functional Representations
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
- Investigating Multi-layer Representations for Dense Passage Retrieval
- AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports
- Fin-ExBERT: User Intent based Text Extraction in Financial Context using Graph-Augmented BERT and trainable Plugin
- Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
- Detecting (Un)answerability in Large Language Models with Linear Directions
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- QoNext: Towards Next-generation QoE for Foundation Models
- A Comprehensive Evaluation of Transformer-Based Question Answering Models and RAG-Enhanced Design
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- CafGa: Customizing Feature Attributions to Explain Language Models
- Concise and Sufficient Sub-Sentence Citations for Retrieval-Augmented Generation
- Integrated Framework for LLM Evaluation with Answer Generation
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- EMO: Pretraining Mixture of Experts for Emergent Modularity
- Just Use XML: Revisiting Joint Translation and Label Projection
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- Riemannian Optimization for LoRA on the Stiefel Manifold
- Identifying and Addressing User-level Security Concerns in Smart Homes Using "Smaller" LLMs
- A Pipeline to Assess Merging Methods via Behavior and Internals
- CompLLM: Compression for Long Context Q&A
- Diversity Boosts AI-Generated Text Detection
- HyperAdapt: Simple High-Rank Adaptation
- False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
- Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages
- SEQR: Secure and Efficient QR-based LoRA Routing
- TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation
- nDNA -- the Semantic Helix of Artificial Cognition
- CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages
- LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- Quantifying Self-Awareness of Knowledge in Large Language Models
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
- Module-Aware Parameter-Efficient Machine Unlearning on Transformers
- SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- InfoGain-RAG: Boosting Retrieval-Augmented Generation via Document Information Gain-based Reranking and Filtering
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- SEDM: Scalable Self-Evolving Distributed Memory for Agents
- Steering MoE LLMs via Expert (De)Activation
- BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
- PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
- Does This Look Familiar to You? Knowledge Analysis via Model Internal Representations
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- HAVE: Head-Adaptive Gating and ValuE Calibration for Hallucination Mitigation in Large Language Models
- HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
- MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages
- Revisiting Third-Party Library Detection: A Ground Truth Dataset and Its Implications Across Security Tasks
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
- Planning with Reasoning using Vision Language World Model
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- DivMerge: A divergence-based model merging method for multi-tasking
- From Confidence to Collapse in LLM Factual Robustness
- MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
- SynDelay: A Synthetic Dataset for Delivery Delay Prediction
- QZhou-Embedding Technical Report
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- QuesGenie: Intelligent Multimodal Question Generation
- Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
- MODE: Mixture of Document Experts for RAG
- Test-time Corpus Feedback: From Retrieval to RAG
- Rethinking Human-Object Interaction Evaluation for both Vision-Language Models and HOI-Specific Methods
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
- VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
- End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
- Identifying and Answering Questions with False Assumptions: An Interpretable Approach
- Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
- Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information
- Wavy Transformer
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models
- Can we Evaluate RAGs with Synthetic Data?
- MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- LaajMeter: A Framework for LaaJ Evaluation
- LyS at SemEval 2025 Task 8: Zero-Shot Code Generation for Tabular QA
- Dynamic Rank Adjustment for Accurate and Efficient Neural Network Training
- AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
- BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
- Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning
- SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
- Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score
- In-Training Defenses against Emergent Misalignment in Language Models
- CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
- Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- KG-Augmented Executable CoT for Mathematical Coding
- PAIRS: Parametric-Verified Adaptive Information Retrieval and Selection for Efficient RAG
- An Entity Linking Agent for Question Answering
- FairLangProc: A Python package for fairness in NLP
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
- SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
- Highlight & Summarize: RAG without the jailbreaks
- Contextually Aware E-Commerce Product Question Answering using RAG
- ProCut: LLM Prompt Compression via Attribution Estimation
- Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
- HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
- WebDS: An End-to-End Benchmark for Web-based Data Science
- D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
- Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
- Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- MUST-RAG: MUSical Text Question Answering with Retrieval Augmented Generation
- Improved Algorithms for Kernel Matrix-Vector Multiplication Under Sparsity Assumptions
- Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- Evaluation and Benchmarking of LLM Agents: A Survey
- MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
- Shapley Uncertainty in Natural Language Generation
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
Related