Lost in the Middle: How Language Models Use Long Contexts
2023/07/06 by Nelson F. Liu, Kevin Lin, Liu, Nelson F. +12 · 10 voices · 1,096 citations
Computer Science · #Advanced Graph Neural Networks #Artificial intelligence #Computer science #Context (archaeology) #History #Key (lock) #Language model #Natural Language Processing Techniques #Natural language processing #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2307.03172
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/07/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Abstract
While recent language models have the ability to take long contexts as input, relatively little is known about how well they use longer context. We analyze the performance of language models on two tasks that require identifying relevant information in their input contexts: multi-document question answering and key-value retrieval. We find that performance can degrade significantly when changing the position of relevant information, indicating that current language models do not robustly make use of information in long input contexts. In particular, we observe that performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models. Our analysis provides a better understanding of how language models use their input context and provides new evaluation protocols for future long-context language models.
Cited by
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
- More Than Memory: Task-Conditioned Signed FFN Writes in Long-Context Retrieval
- Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- Lost in Context: Addressing Context Anxiety in Large Language Models
- Simulating Human Memory with Language Models
- State Compression in Two-Agent LLM Relays: A Closed-World Study of Constraint Preservation
- Accurate and Efficient Long-Term Memory for LLM Agents
- RE-AD: Real-Time Requirement Adherence for Data Labeling
- Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis
- Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
- Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval
- LLMs Corrupt Your Documents When You Delegate
- Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
- Position: Modular Memory is the Key to Continual Learning Agents
- Mitigating Conversational Inertia in Multi-Turn Agents
- Do LLMs Benefit From Their Own Words?
- Perplexity Cannot Always Tell Right from Wrong
- Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- Accumulating Context Changes the Beliefs of Language Models
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- The Impossibility of Inverse Permutation Learning in Transformer Models
- What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
- EntropyLong: Effective Long-Context Training via Predictive Uncertainty
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks
- GAME: Genomic API for Model Evaluation
- JavelinGuard: Low-Cost Transformer Architectures for LLM Security
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- LLMs Get Lost In Multi-Turn Conversation
- Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers.
- Enough Coin Flips Can Make LLMs Act Bayesian
- Multi-Token Attention
- ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
- Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
- General Intelligence Requires Reward-based Pretraining
- DMA: Online RAG Alignment with Human Feedback
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks
- Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
- MatKV: Trading Compute for Flash Storage in LLM Inference
- From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
- MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents
- Synthesizing Multi-Agent Harnesses for Vulnerability Discovery
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Towards a Relevance Posterior in Neural Information Access
- ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Accelerating Language Model Workflows with Prompt Choreography
- ACM: Agentic Context Management for Long Horizon Tasks
- Generative AI for Requirements Engineering: A Systematic Literature Review
- TARGET: Automated Scenario Generation from Traffic Rules for Testing Autonomous Vehicles via Validated LLM-Guided Knowledge Extraction
- A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
- KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
- VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy
- FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
- LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
- Specula: Scaling formal specifications for autonomous model checking of system code
- COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
- Ranked by Position: Order Sensitivity as an Exploitable Attack Surface in LLM Listwise Recommenders
- How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
- When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- Let AI Agents Translate Networks, Not Reason About Them
- CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents
- RoboMME-Interference: Benchmarking Robot Memory Under Interference
- AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System
- From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance
- CRAFT: Learn the Schema, Execute the Plan
- Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
- Three Sides of Retrieval: Factorial Evidence for Document-Side, Query-Side, and Answer-Side Complementarity in RAG
- Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations
- Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models
- The Cartesian Cut in Agentic AI
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation
- Beyond Context: Large Language Models' Failure to Grasp Users' Intent
- Making AI Functional with Workarounds: An Insider's Account of Invisible Labour in Organisational Politics
- Assessing the Software Security Comprehension of Large Language Models
- ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- Decoupling Positional and Symbolic Attention Behavior in Transformers
- Learning to Wait: Synchronizing Agents with the Physical World
- The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
- Integrating Large Language Models and Knowledge Graphs to Capture Political Viewpoints in News Media
- Dynamic Context Selection for Retrieval-Augmented Generation: Mitigating Distractors and Positional Bias
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Context-Picker: Dynamic context selection using multi-stage reinforcement learning
- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
- Uncovering the Role of Initial Saliency in U-Shaped Attention Bias: Scaling Initial Token Weight for Enhanced Long-Text Processing
- Context Branching for LLM Conversations: A Version Control Approach to Exploratory Programming
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation
- CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
- Automated Penetration Testing with LLM Agents and Classical Planning
- KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration
- Replace, Don't Expand: Mitigating Context Dilution in Multi-Hop RAG via Fixed-Budget Evidence Assembly
- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
- Workflow is All You Need: Escaping the "Statistical Smoothing Trap" via High-Entropy Information Foraging and Adversarial Pacing
- Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology
- Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems
- An Index-based Approach for Efficient and Effective Web Content Extraction
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- Modeling Contextual Passage Utility for Multihop Question Answering
- Efficient Text Classification with Conformal In-Context Learning
- Sift or Get Off the PoC: Applying Information Retrieval to Vulnerability Research with SiftRank
- MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
- Intrinsically Interpretable Attention via Sparse Post-Training
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
- The Road of Adaptive AI for Precision in Cybersecurity
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms
- HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
- Leveraging LLMs for Structured Data Extraction from Unstructured Patient Records
- Input Order Shapes LLM Semantic Alignment in Multi-Document Summarization
- When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
- Probing the "Psyche'' of Large Reasoning Models: Understanding Through a Human Lens
- MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
- Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
- Mitigating Semantic Drift: Evaluating LLMs' Efficacy in Psychotherapy through MI Dialogue Summarization
- Auditable Context-Aware HFMD Forecasting with Structured LLM Agents
- Masks Can Be Distracting: On Context Comprehension in Diffusion Language Models
- SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
- Translating Large-Scale C Repositories to Idiomatic Rust
- InvisibleBench: A Deployment Gate for Caregiving Relationship AI
- Large Language Model Aided Birt-Hogg-Dube Syndrome Diagnosis with Multimodal Retrieval-Augmented Generation
- Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
- Training-Free Active Learning Framework in Materials Science with Large Language Models
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- SmartPoC: Generating Executable and Validated PoCs for Smart Contract Bug Reports
- Enhancing Large Language Models for Automated Homework Assessment in Undergraduate Circuit Analysis
- Rethinking Retrieval: From Traditional Retrieval Augmented Generation to Agentic and Non-Vector Reasoning Systems in the Financial Domain for Large Language Models
- ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
- Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
- A Simple Yet Strong Baseline for Long-Term Conversational Memory of LLM Agents
- Learning to Debug: LLM-Organized Knowledge Trees for Solving RTL Assertion Failures
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
- CARE-RAG - Clinical Assessment and Reasoning in RAG
- Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- A Viable Paradigm of Software Automation: Iterative End-to-End Automated Software Development
- From Topology to Behavioral Semantics: Enhancing BGP Security by Understanding BGP's Language with LLMs
- SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
- Hint-Augmented Re-ranking: Efficient Product Search using LLM-Based Query Decomposition
- BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
- DualGR: Generative Retrieval with Long and Short-Term Interests Modeling
- Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
- A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
- Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
- ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
- A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- Chain of Summaries: Summarization Through Iterative Questioning
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- Stress Testing Factual Consistency Metrics for Long-Document Summarization
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
- TreeWriter: AI-Assisted Hierarchical Planning and Writing for Long-Form Documents
- Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
- In the beginning was the Word: LLM-VaR and LLM-ES
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- Discourse Graph Guided Document Translation with Large Language Models
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Retrieval Quality at Context Limit
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- Large Language Models for Explainable Threat Intelligence
- Cambrian-S: Towards Spatial Supersensing in Video
- RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
- LLM-enhanced Air Quality Monitoring Interface via Model Context Protocol
- Test-Time Adaptation for LLM Agents via Environment Interaction
- Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
- GEMMA-SQL: A Novel Text-to-SQL Model Based on Large Language Models
- LGM: Enhancing Large Language Models with Conceptual Meta-Relations and Iterative Retrieval
- Jailbreaking in the Haystack
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- Using Span Queries to Optimize for Cache and Attention Locality
- KV Cache Transform Coding for Compact Storage in LLM Inference
- Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- How Focused Are LLMs? A Quantitative Study via Repetitive Deterministic Prediction Tasks
- ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
- DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
- Glia: A Human-Inspired AI for Automated Systems Design and Optimization
- Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline
- Predicate Renaming via Large Language Models
- AutoSurvey2: Empowering Researchers with Next Level Automated Literature Surveys
- Beyond Long Context: When Semantics Matter More than Tokens
- User Misconceptions of LLM-Based Conversational Programming Assistants
- Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
- Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- A Graph-Native Bitemporal Memory Store for Conversational AI Agents
- CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- The New Shape of Search: How Conversational AI Recomposes Information Seeking
- GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations
- LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability
- Attention-Aligned Reasoning for Large Language Models
- Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising
- Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimization
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
- POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
- SPIRE: Structure-Preserving Interpretable Retrieval of Evidence
- Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion
- Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
- Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
- EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
- RGMem: Renormalization Group-inspired Memory Evolution for Language Agents
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
- StorageXTuner: An LLM Agent-Driven Automatic Tuning Framework for Heterogeneous Storage Systems
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems
- Structured Interfaces for Automated Reasoning with 3D Scene Graphs
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- What do vision-language models see in the context? Investigating multimodal in-context learning
- ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents
- Evaluating Long-Term Memory for Long-Context Question Answering
- COOPERA: Continual Open-Ended Human-Robot Assistance
- QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
- CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
- Tagging-Augmented Generation: Assisting Language Models in Finding Intricate Knowledge In Long Contexts
- Leveraging Large Language Models to Identify Conversation Threads in Collaborative Learning
- Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- ATOM: AdapTive and OptiMized dynamic temporal knowledge graph construction using LLMs
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
- FAIR-RAG: Faithful Adaptive Iterative Refinement for Retrieval-Augmented Generation
- Deep Literature Survey Automation with an Iterative Workflow
- Redefining Retrieval Evaluation in the Era of LLMs
- Paper2Web: Let's Make Your Paper Alive!
- Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
- Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
- Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- LM-mixup: Text Data Augmentation via Language Model based Mixup
- ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
- Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
- CMT-Bench: Cricket Multi-Table Generation Benchmark for Probing Robustness in Large Language Models
- CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows
- Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!
- AcademicEval: Live Long-Context LLM Benchmark
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- TaxoAlign: Scholarly Taxonomy Generation Using Language Models
- StreamingThinker: Large Language Models Can Think While Reading
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
- More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
- Pursuing Minimal Sufficiency in Spatial Reasoning
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- Boosting Instruction Following at Scale
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
- Rethinking Schema Linking: A Context-Aware Bidirectional Retrieval Approach for Text-to-SQL
- PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
- Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
- Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models
- Toward Cybersecurity-Expert Small Language Models
- BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
- MADREC: A Multi-Aspect Driven LLM Agent for Explainable and Adaptive Recommendation
- Document Intelligence in the Era of Large Language Models: A Survey
- D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
- Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
- LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
- MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
- Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
- Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
- Not in Sync: Unveiling Temporal Bias in Audio Chat Models
- DSAS: A Universal Plug-and-Play Framework for Attention Optimization in Multi-Document Question Answering
- Are Large Language Models Effective Knowledge Graph Constructors?
- The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
- Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling
- Scaling Long-Horizon LLM Agent via Context-Folding
- Glance for Context: Learning When to Leverage LLMs for Node-Aware GNN-LLM Fusion
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Traj-CoA: Patient Trajectory Modeling via Chain-of-Agents for Lung Cancer Risk Prediction
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- Style Over Story: A Process-Oriented Study of Authorial Creativity in Large Language Models
- Doc-to-Atom: Learning to Compile and Compose Memory Atoms
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- Lost in the Middle: An Emergent Property from Information Retrieval Demands in LLMs
- How Good Are LLMs at Processing Tool Outputs?
- PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection
- GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
- LLP: LLM-based Product Pricing in E-commerce
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
- KORMo: Korean Open Reasoning Model for Everyone
- Cluster-based Adaptive Retrieval: Dynamic Context Selection for RAG Applications
- A Locally Executable AI System for Improving Preoperative Patient Communication: A Multi-Domain Clinical Evaluation
- Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- LLM-Assisted Web Measurements
- Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity
- Robust Heuristic Algorithm Design with LLMs
- Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
- CaRT: Teaching LLM Agents to Know When They Know Enough
- ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
- Traceability and Accountability in Role-Specialized Multi-Agent LLM Pipelines
- In-Context Clustering with Large Language Models
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- Leveraging LLMs to Streamline the Review of Public Funding Applications
- Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
- Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- Beyond More Context: How Granularity and Order Drive Code Completion Quality
- Extending ResourceLink: Patterns for Large Dataset Processing in MCP Applications
- LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection
- On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
- Revisiting Long-context Modeling from Context Denoising Perspective
- Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
- Constraint-Aware Route Recommendation from Natural Language via Hierarchical LLM Agents
- Prototype-Based Dynamic Steering for Large Language Models
- Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- Slm-mux: Orchestrating small language models for reasoning
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- COSMIR: Chain Orchestrated Structured Memory for Iterative Reasoning over Long Context
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
- Exploring Network-Knowledge Graph Duality: A Case Study in Agentic Supply Chain Risk Analysis
- CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
- Hearing the Order: Investigating Selection Bias in Large Audio-Language Models
- Memory-Augmented Log Analysis with Phi-4-mini: Enhancing Threat Detection in Structured Security Logs
- LongCodeZip: Compress Long Context for Code Language Models
- TokMem: Tokenized Procedural Memory for Large Language Models
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
- AuditAgent: Expert-Guided Multi-Agent Reasoning for Cross-Document Fraudulent Evidence Discovery
- CreAgentive: An Agent Workflow Driven Multi-Category Creative Generation Engine
- AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
- Interactive Learning for LLM Reasoning
- QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
- SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
- End-to-End Aspect-Guided Review Summarization at Scale
- RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
- A Field Guide to Deploying AI Agents in Clinical Practice
- Detecting and Fixing API Misuses of Data Science Libraries Using Large Language Models
- SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
- Comparing Open-Source and Commercial LLMs for Domain-Specific Analysis and Reporting: Software Engineering Challenges and Design Trade-offs
- TENET: Leveraging Tests Beyond Validation for Code Generation
- SynthPert: Enhancing LLM Biological Reasoning via Synthetic Reasoning Traces for Cellular Perturbation Prediction
- Rethinking and Benchmarking Large Language Models for Graph Reasoning
- Expanding Computation Spaces of LLMs at Inference Time
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- MMPB: It's Time for Multi-Modal Personalization
- Context Parametrization with Compositional Adapters
- Review of Hallucination Understanding in Large Language and Vision Models
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding
- Graph of Agents: Principled Long Context Modeling by Emergent Multi-Agent Collaboration
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- Same Content, Different Representations: A Controlled Study for Table QA
- In-Context Learning can Perform Continual Learning Like Humans
- Binary Autoencoder for Mechanistic Interpretability of Large Language Models
- Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
- RepLLM: Toward Automatically Reproducing Network Research Results
- Efficient On-Device Agents via Adaptive Context Management
- PromptDebt: A Comprehensive Study of Technical Debt Across LLM Projects
- DRES: Benchmarking LLMs for Disfluency Removal
- A co-evolving agentic AI system for medical imaging analysis
- MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection
- RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
- Mamba Modulation: On the Length Generalization of Mamba
- Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
- Reverse Engineering User Stories from Code using Large Language Models
- Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- SKIMIX: Multi-Agent Harness-Time Scaling with Skill Mixture for Dynamic Harness Engineering
- The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials
- RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
- Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
- Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
- Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories
- OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing
- How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
- Inspectable AI for Science: A Research Object Approach to Generative AI Governance
- Non-Parametric Structural Priors for Geometry Theorem Prediction
- PoRe: Position-Reweighted Visual Token Pruning for Vision Language Models
- Automated Extraction of Material Properties using LLM-based AI Agents
- CompLLM: Compression for Long Context Q&A
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Anecdoctoring: Automated Red-Teaming Across Language and Place
- LingoQ: Bridging the Gap between ESL Learning and Work through AI-Generated Work-Related Quizzes
- XaaS Containers: Performance-Portable Representation With Source and IR Containers
- Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
- AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
- How Persuasive is Your Context?
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments
- nDNA -- the Semantic Helix of Artificial Cognition
- Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
- Influence Guided Context Selection for Effective Retrieval-Augmented Generation
- EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
- ST-Raptor: LLM-Powered Semi-Structured Table Question Answering
- Relevance to Utility: Process-Supervised Rewrite for RAG
- TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
- Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
- A Taxonomy of Prompt Defects in LLM Systems
- Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
- Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
- DSPC: Dual-Stage Progressive Compression Framework for Efficient Long-Context Reasoning
- Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
- An Empirical Study on Failures in Automated Issue Solving
- Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
- Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Less Is More: Elevating RAG via Performance-Driven Context Compression
- From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
- The Few-shot Dilemma: Over-prompting Large Language Models
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- ConvergeWriter: Data-Driven Bottom-Up Article Construction
- Token Homogenization under Positional Bias
- FinGEAR: Financial Mapping-Guided Enhanced Answer Retrieval
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- Context-Aware Language Models for Forecasting Market Impact from Sequences of Financial News
- UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
- Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary Adjudication
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Beyond Token Limits: Assessing Language Model Performance on Long Text Classification
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
- Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
- ObjexMT: Objective Extraction and Metacognitive Calibration for LLM-as-a-Judge under Multi-Turn Jailbreaks
- Deploying AI for Signal Processing education: Selected challenges and intriguing opportunities
- Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation
- Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
- Multi-view-guided Passage Reranking with Large Language Models
- Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
- Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning
- Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
- Language Native Lightly Structured Databases for Large Language Model Driven Composite Materials Research
- GRACE: Graph-Guided Repository-Aware Code Completion through Hierarchical Code Fusion
- X-SQL: Expert Schema Linking and Understanding of Text-to-SQL with Multi-LLMs
- Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
- LLM-as-classifier: Semi-Supervised, Iterative Framework for Hierarchical Text Classification using Large Language Models
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- CMR-SPB: Cross-Modal Multi-Hop Reasoning over Text, Image, and Speech with Path Balance
- Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
- LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
- LLM-Assisted Semantic Alignment and Integration in Collaborative Model-Based Systems Engineering Using SysML v2
- Generative Goal Modeling
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- TableZoomer: A Collaborative Agent Framework for Large-scale Table Question Answering
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- Can Multi-turn Self-refined Single Agent LMs with Retrieval Solve Hard Coding Problems?
- Memory Limitations of Prompt Tuning in Transformers
- Access Paths for Efficient Ordering with Large Language Models
- OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Addressing accuracy and hallucination of LLMs in Alzheimer's disease research through knowledge graphs
- OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models
- An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- Joint Enhancement of Relational Reasoning for Long-Context LLMs
- Research Challenges in Relational Database Management Systems for LLM Queries
- GDS Agent for Graph Algorithmic Reasoning
- Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
- TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
- PG-Agent: An Agent Powered by Page Graph
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
- Robustness is Important: Limitations of LLMs for Data Fitting
- Orchid: Orchestrating Context Across Creative Workflows with Generative AI
- An Investigation on Group Query Hallucination Attacks
- Test-time Corpus Feedback: From Retrieval to RAG
- Text to Query Plans for Question Answering on Large Tables
- REALM: Recursive Relevance Modeling for LLM-based Document Re-Ranking
- Agri-Query: A Case Study on RAG vs. Long-Context LLMs for Cross-Lingual Technical Question Answering
- Named Entity Recognition of Historical Text via Large Language Model
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
- Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- User-Assistant Bias in LLMs
- Benchmarking LLM-based agents for single-cell omics analysis
- SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communication
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- RAG for Geoscience: What We Expect, Gaps and Opportunities
- ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning
- PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
- Intent-Aware Schema Generation And Refinement For Literature Review Tables
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- APIO: Automatic Prompt Induction and Optimization for Grammatical Error Correction and Text Simplification
- Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
- Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory
- SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
- Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
- Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Positional Biases Shift as Inputs Approach Context Window Limits
- Triple-S: A Collaborative Multi-LLM Framework for Solving Long-Horizon Implicative Tasks in Robotics
- Vec2Summ: Text Summarization via Probabilistic Sentence Embeddings
- A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks
- Synthesizing scientific literature with retrieval-augmented language models
- Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- Leveraging LLMs for Privacy-Aware Predictions in Participatory Budgeting
- Attention Basin: Why Contextual Position Matters in Large Language Models
- BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation
- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
- TURA: Tool-Augmented Unified Retrieval Agent for AI Search
- Small transformer architectures for task switching
- Characterizing Deep Research: A Benchmark and Formal Definition
- A Pragmatist Robot: Learning to Plan Tasks by Experiencing the Real World
- MOTIF: Multi-strategy Optimization via Turn-based Interactive Framework
- Semantic-aware Graph-guided Behavior Sequences Generation with Large Language Models for Smart Homes
- Write, Rank, or Rate: Comparing Methods for Studying Visualization Affordances
- Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science
- Key-Augmented Neural Triggers for Knowledge Sharing
- Do language models accommodate their users? A study of linguistic convergence
- Long Story Generation via Knowledge Graph and Literary Theory
- Cognitive Loop via In-Situ Optimization: Self-Adaptive Reasoning for Science
- Web3 x AI Agents: Landscape, Integrations, and Foundational Challenges
- LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
- Trainable Dynamic Mask Sparse Attention
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- Evaluating Position Bias in Large Language Model Recommendations
- ProCut: LLM Prompt Compression via Attribution Estimation
- ReflecSched: Solving Dynamic Flexible Job-Shop Scheduling via LLM-Powered Hierarchical Reflection
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- Court of LLMs: Evidence-Augmented Generation via Multi-LLM Collaboration for Text-Attributed Graph Anomaly Detection
- GETALP@AutoMin 2025: Leveraging RAG to Answer Questions based on Meeting Transcripts
- Enhancing Retrieval-Augmented Generation for Electric Power Industry Customer Support
- ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
- NeedleChain: Measuring Intact Long-Context Reasoning Capability of Large Language Models
- Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation
- DGP: A Dual-Granularity Prompting Framework for Fraud Detection with Graph-Enhanced LLMs
- A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations
- FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression
- Prometheus: Towards Long-Horizon Codebase Navigation for Repository-Level Problem Solving
- Reading Between the Timelines: RAG for Answering Diachronic Questions
- Flora: Effortless Context Construction to Arbitrary Length and Scale
- CaliDrop: KV Cache Compression with Calibration
- SelfRACG: Enabling LLMs to Self-Express and Retrieve for Code Generation
- A Survey of Multimodal Hallucination Evaluation and Detection
- Input Reduction Enhanced LLM-based Program Repair
- Can LLMs Generate User Stories and Assess Their Quality?
- StyleAdaptedLM: Enhancing Instruction Following Models with Efficient Stylistic Transfer
- MemoCoder: Automated Function Synthesis using LLM-Supported Agents
- PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
- Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
- Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
- Token Reduction Is Not Cost Reduction
- Rethinking MCP Security: A Large-Scale Study of Runtime MCP Servers and Security Scanner Reliability
- PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents
- Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs
- Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
- Eywa: Provenance-Grounded Long-Term Memory for AI Agents
- MeMo: Memory as a Model
- Autonomous LLM Agent Worms: Cross-Platform Propagation, Automated Discovery and Temporal Re-Entry Defense
- Natural-Language Agent Harnesses
- Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias
- IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering
- Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
- Exploiting Primacy Effect To Improve Large Language Models
- SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
- Agent Identity Evals: Measuring Agentic Identity
- Enhancing patent retrieval using automated patent summarization
- A Unifying Scheme for Extractive Content Selection Tasks
- Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
- Learning to Reason Across Parallel Samples for LLM Reasoning
- RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
- Formal Methods Meets Readability: Auto-Documenting JML Java Code
- Exploring the use of LLMs in the Italian legal domain: A survey on recent applications
- From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection
- AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos
- Journalism-Guided Agentic In-Context Learning for News Stance Detection
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
- DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
- Mitigating Posterior Salience Attenuation in Long-Context LLMs with Positional Contrastive Decoding
- Am I on the Right Track? What Can Predicted Query Performance Tell Us about the Search Behaviour of Agentic RAG
- Draft-based Approximate Inference for LLMs
- Retrieval-Augmented Recommendation Explanation Generation with Hierarchical Aggregation
- GraphRunner: A Multi-Stage Framework for Efficient and Accurate Graph-Based Retrieval
- Structure-Augmented Reasoning Generation
- InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching
- FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-based RAG via Multi-Partite Graph and Query-Driven Iterative Retrieval
- Know It, Act on It: Investigating Memory Utilization in LLM Personalization
- DrugMCTS: a drug repurposing framework combining multi-agent, RAG and Monte Carlo Tree Search
- Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies
- Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
- RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
- TransMem: Transforming Hidden States into Memory for Large Language Models
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
- Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering
- KERAGR: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation
- Self-Review Framework for Enhancing Instruction Following Capability of LLM
- PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
- Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
- "Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models
- XiYan-SQL: A Novel Multi-Generator Framework For Text-to-SQL
- Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
- LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
- An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
- Can Large Language Models Automate the Refinement of Cellular Network Specifications?
- LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers
- SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
- BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering
- Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering
- CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate
- MemOS: A Memory OS for AI System
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs
- Legal Requirements Translation from Law
- ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
- Quantifying Cognitive Bias Induction in LLM-Generated Content
- CROP: Circuit Retrieval and Optimization with Parameter Guidance using LLMs
- The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems
- Decomposing Prediction Mechanisms for In-Context Recall
- Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
- Dynamic and Parametric Retrieval-Augmented Generation
- Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Holistic Artificial Intelligence in Medicine; improved performance and explainability
- RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence
- Evaluating and Improving Large Language Models for Competitive Program Generation
- JointRank: Rank Large Set with Single Pass
- Lost at the Beginning of Reasoning
- UiS-IAI@LiveRAG: Retrieval-Augmented Information Nugget-Based Generation of Responses
- Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation
- Small Encoders Can Rival Large Decoders in Detecting Groundedness
- Ad-Hoc Human-AI Coordination Challenge
- π-CoT: Prolog-Initialized Chain-of-Thought Prompting for Multi-Hop Question-Answering
- Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
- KnowML: Improving Generalization of ML-NIDS with Attack Knowledge Graphs
- PEVLM: Parallel Encoding for Vision-Language Models
- Unable to Forget: Proactive Interference Reveals Working Memory Limits in LLMs Beyond Context Length
- heiDS at ArchEHR-QA 2025: From Fixed-k to Query-dependent-k for Retrieval Augmented Generation
- Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation
- Controlled Retrieval-augmented Context Evaluation for Long-form RAG
- FORGE: An LLM-driven Framework for Large-Scale Smart Contract Vulnerability Dataset Construction
- Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models
- Use Property-Based Testing to Bridge LLM Code Generation and Validation
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- Markov-Enhanced Clustering for Long Document Summarization: Tackling the 'Lost in the Middle' Challenge with Large Language Models
- TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
- Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
- Sequence-to-Sequence Models with Attention Mechanistically Map to the Architecture of Human Memory Search
- Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
- Evaluating the Use of LLMs for Documentation to Code Traceability
- SGIC: A Self-Guided Iterative Calibration Framework for RAG
- Reranking-based Generation for Unbiased Perspective Summarization
- StoryWriter: A Multi-Agent Framework for Long Story Generation
- Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
- The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Statistical Hypothesis Testing for Auditing Robustness in Language Models
- TokenShapley: Token Level Context Attribution with Shapley Value
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
- Don't throw the baby out with the bathwater: How and why deep learning for ARC
- StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework
- InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
- What Makes a Good Natural Language Prompt?
- GenerationPrograms: Fine-grained Attribution with Executable Programs
- Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
- Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
- GeometryZero: Improving Geometry Solving for LLM with Group Contrastive Policy Optimization
- Leveraging In-Context Learning for Language Model Agents
- Scaling Algorithm Distillation for Continuous Control with Mamba
- Atomic Reasoning for Scientific Table Claim Verification
- TermSight: Making Service Contracts Approachable
- Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
- LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
- Post Persona Alignment for Multi-Session Dialogue Generation
- Long-Short Alignment for Effective Long-Context Modeling in LLMs
- Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models
- No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
- Automatic Speech Recognition of African American English: Lexical and Contextual Effects
- Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
- United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory
- Beyond Facts: Evaluating Intent Hallucination in Large Language Models
- Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge
- BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions
- Respecting Temporal-Causal Consistency: Entity-Event Knowledge Graphs for Retrieval-Augmented Generation
- Can Theoretical Physics Research Benefit from Language Agents?
- CoMemo: LVLMs Need Image Context with Image Memory
- Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment
- Cartridges: Lightweight and general-purpose long context representations via self-study
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- Dynamic Context Tuning for Retrieval-Augmented Generation: Enhancing Multi-Turn Planning and Tool Adaptation
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning
- Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement
- Context Is Not Comprehension
- ECoRAG: Evidentiality-guided Compression for Long Context RAG
- Inference-Time Hyper-Scaling with KV Cache Compression
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
- CDE-Mapper: Using Retrieval-Augmented Language Models for Linking Clinical Data Elements to Controlled Vocabularies
- Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
- Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL
- Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
- KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
- Enhancing Large Language Models with Neurosymbolic Reasoning for Multilingual Tasks
- An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models
- A Controllable Examination for Long-Context Language Models
- Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering
- Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
- Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models
- Scaling Textual Gradients via Sampling-Based Momentum
- TreeRare: Syntax Tree-Guided Retrieval and Reasoning for Knowledge-Intensive Question Answering
- Inter-Passage Verification for Multi-evidence Multi-answer QA
- Can Large Language Models Predict Parallel Code Performance?
- Harnessing Large Language Models for Scientific Novelty Detection
- NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models
- REOrdering Patches Improves Vision Models
- Sentinel: Attention Probing of Proxy Models for LLM Context Compression with an Understanding Perspective
- LogiDebrief: A Signal-Temporal Logic based Automated Debriefing Approach with Large Language Models Integration
- BLAB: Brutally Long Audio Bench
- ARC: Argument Representation and Coverage Analysis for Zero-Shot Long Document Summarization with Instruction Following LLMs
- How Does Response Length Affect Long-Form Factuality
- mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
- ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
- Curse of High Dimensionality Issue in Transformer for Long-context Modeling
- MapStory: Prototyping Editable Map Animations with LLM Agents
- Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
- Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations
- The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
- Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
- RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
- CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models
- Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
- Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
- Can Past Experience Accelerate LLM Reasoning?
- Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
- VeriTrail: Closed-Domain Hallucination Detection with Traceability
- SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
- Code Researcher: Deep Research Agent for Large Systems Code and Commit History
- Effectiveness of Prompt Optimization in NL2SQL Systems
- MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents
- Two Causally Related Needles in a Video Haystack
- Token-Importance Guided Direct Preference Optimization
- Select, Read, and Write: A Multi-Agent Framework of Full-Text-based Related Work Generation
- The Role of Diversity in In-Context Learning for Large Language Models
- DocMEdit: Towards Document-Level Model Editing
- From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- Direct Retrieval-augmented Optimization: Synergizing Knowledge Selection and Language Models
- Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents
- A Survey on Progress in LLM Alignment from the Perspective of Reward Design
- LLMs as Better Recommenders with Natural Language Collaborative Signals: A Self-Assessing Retrieval Approach
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
- SQUiD: Synthesizing Relational Databases from Unstructured Text
- Weaver: Interweaving SQL and LLM for Table Reasoning
- 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
- Language Models Should be Used to Surface the Unwritten Code of Science and Society
- GenAI and the Mirage of Personalised Learning for All
- Hypercube-Based Retrieval-Augmented Generation for Scientific Question-Answering
- Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation
- TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory
- Knoll: Creating a Knowledge Ecosystem for Large Language Models
- Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
- QwenLong-CPRS: Towards ∞-LLMs with Dynamic Context Optimization
- Towards Practical Defect-Focused Automated Code Review
- Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour
- Self-Improving Large Language Models via Progressive Experience Evolution
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
- PD3: A Project Duplication Detection Framework via Adapted Multi-Agent Debate
- Learning to Focus: Context Extraction for Efficient Code Vulnerability Detection with Language Models
- GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations
- MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
- Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
- CUB: Benchmarking Context Utilisation Techniques for Language Models
- MuseScorer: Idea Originality Scoring At Scale
- Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
- From Compression to Expression: A Layerwise Analysis of In-Context Learning
- Understanding Differential Transformer Unchains Pretrained Self-Attentions
- MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
- Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
- Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
- Revealing Language Model Trajectories via Kullback-Leibler Divergence
- P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
- Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
- Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
- Set-LLM: A Permutation-Invariant LLM
- IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
- Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
- Do RAG Systems Really Suffer From Positional Bias?
- In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis
- Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning
- Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
- An Empirical Study of Position Bias in Modern Information Retrieval
- Table Foundation Models: on knowledge pre-training for tabular learning
- NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
- Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
- MAFA: A multi-agent framework for annotation
- Transparent and Robust RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability
- PromptPrism: A Linguistically-Inspired Taxonomy for Prompts
- Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Understanding
- The Hidden Dangers of Browsing AI Agents
- EAVIT: Efficient and Accurate Human Value Identification from Text data via LLMs
- CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
- AMAQA: A Metadata-based QA Dataset for RAG Systems
- Batched Self-Consistency Improves LLM Relevance Assessment and Ranking
- V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
- ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
- Automated Profile Inference with Language Model Agents
- G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment
- Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation
- Do Code LLMs Do Static Analysis?
- Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
- Context Compaction Theory
- Talk to Your Slides: Language-Driven Agents for Efficient Slide Editing
- The Epistemic Politics of AI Anthropomorphism
- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents
- An AI Approach to Verified Production Cryptographic Libraries
- RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation
- Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture Slides
- GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation
- Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
- Accurate KV Cache Quantization with Outlier Tokens Tracing
- Creating General User Models from Computer Use
- A Systematic Analysis of Base Model Choice for Reward Modeling
- Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
- THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering
- GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
- SafeTrans: LLM-assisted Transpilation from C to Rust
- The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
- Safety Invariants for Agents Orchestrating Irreversible State Transitions: A Four-Dimensional Formalism Evaluated on Public Ledgers
- Demystifying AI Agents: The Final Generation of Intelligence
- AC-LoRA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs
- LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
- SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering
- Reasoning Capabilities of Large Language Models on Dynamic Tasks
- CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability
- Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
- Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation
- Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies
- LLM-based Prompt Ensemble for Reliable Medical Entity Recognition from EHRs
- Answer Presence Drives RAG Rewriting Gains
- AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents
- Putting It All into Context: Simplifying Agents with LCLMs
- Overflow Prevention Enhances Long-Context Recurrent LLMs
- GRADA: Graph-based Reranking against Adversarial Documents Attack
- The Distracting Effect: Understanding Irrelevant Passages in RAG
- OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
- MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG
- From Millions of Tweets to Actionable Insights: Leveraging LLMs for User Profiling
- Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
- ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents
- Automating Database-Native Function Code Synthesis with LLMs
- Contextual Drag: How Errors in the Context Affect LLM Reasoning
- In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts
- Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
- CodeSSM: Towards State Space Models for Code Understanding
- Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off
- MerLean-Prover: A Recursive Looping Harness for Lean 4 Theorem Proving
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- Classifier Context Rot: Monitor Performance Degrades with Context Length
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
- LLM Agents in Law: Taxonomy, Applications, and Challenges
- Improving Coherence and Persistence in Agentic AI for System Optimization
- Teaching an Agent to Sketch One Part at a Time
- The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
- OrchestrXR: A Multi-Agent System for Idea-to-Prototype XR Study Authoring
- Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
- Diagnosing and Mitigating Context Rot in Long-horizon Search
- D-MEM: Dopamine-Gated Agentic Memory via Reward Prediction Error Routing
- Aeon: High-Performance Neuro-Symbolic Memory Management for Long-Horizon LLM Agents
- STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
- Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation
- Learning the ARTS of Search for Automated Discovery
- Large Language Models Do Not Always Need Readable Language
- AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Facts
- Structured Inference with Large Language Gibbs
- SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG
- Uncertainty-Aware Hybrid Retrieval for Long-Document RAG
- Draft-Refine-Optimize: Self-Evolved Learning for Natural Language to MongoDB Query Generation
- The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
- The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
- Does My README File Need To Be Updated? Exploring LLM-Based README Maintenance
- Evolutionary Context Search for Automated Skill Acquisition
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
- Panini: Continual Learning in Token Space via Structured Memory
- Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models
- Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection
- What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
- SkillsInjector: Dynamic Skill Context Construction for LLM Agents
- Semantic Chunking and the Entropy of Natural Language
- Using Large Language Models in Physics Education
- Parallel Context Compaction for Long-Horizon LLM Agent Serving
- ASSEMBLAGE-DEEPHISTORY: A Cross-Build Binary Dataset with Temporal Coverage
- LongFuncEval: Measuring the effectiveness of long context models for function calling
- Harnessing AtomisticSkills for Agentic Atomistic Research
- APWA: A Distributed Architecture for Parallelizable Agentic Workflows
- Safety and accuracy follow different scaling laws in clinical large language models
- State Representation and Termination for Recursive Reasoning Systems
- MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
- An Agentic Approach to Metadata Reasoning
- Transforming Evidence Synthesis: A Systematic Review of the Evolution of Automated Meta-Analysis in the Age of AI
- Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
- Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
- Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
- A Cloud-Native Architecture for Human-in-Control LLM-Assisted OpenSearch in Investigative Settings
- Learning Agent-Compatible Context Management for Long-Horizon Tasks
- Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
- Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems
- Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft Game
- GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
- Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
- Screening Is Enough
- Coding Agents are Effective Long-Context Processors
- Facts as First Class Objects: Knowledge Objects for Persistent LLM Memory
- Interpretable Context Methodology: Folder Structure as Agentic Architecture
- A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology
- GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
- NanoKnow: How to Know What Your Language Model Knows
- Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
- Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
- Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
- Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
- SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
- Towards Large Language Models for Lunar Mission Planning and In Situ Resource Utilization
- Think Before Writing: Feature-Level Multi-Objective Optimization for Generative Citation Visibility
- Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
- Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)
- Federation over Text: Insight Sharing for Multi-Agent Reasoning
- Generative Product Recommendations for Implicit Superlative Queries
- HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
- Application and Optimization of Large Models Based on Prompt Tuning for Fact-Check-Worthiness Estimation
- AgentSPEX: An Agent SPecification and EXecution Language
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context
- Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
- Large Language Models for Departmental Expert Review Quality Scores
- Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
- EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
- You Need Better Attention Priors
- Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
- AI Agents Need Memory Control Over More Context
- Position Bias Undermines Preference Consistency in Listwise LLM-Based Reranking
- DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
- SimpleMem: Efficient Lifelong Memory for LLM Agents
- Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
- Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
- An ontology-based retrieval augmented generation procedure for a voice-controlled maintenance assistant
- DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question Answering
- MODP: Multi Objective Directional Prompting
- EvidenceBench: A Benchmark for Extracting Evidence from Biomedical Papers
- Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
- LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension
- An Empirical Study on Prompt Compression for Large Language Models
- LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams
- A RAG-Based Multi-Agent LLM System for Natural Hazard Resilience and Adaptation
- In-Context Learning can distort the relationship between sequence likelihoods and biological fitness
- Credible Plan-Driven RAG Method for Multi-Hop Question Answering
- Generative Optimization for Incentivized Advertising with Global Level Constraints
- FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory
- Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression
- PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates
- MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off
- Neighborhood-Aware Dual Biomedical Entity Linking
- The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents
- Contextual Agentic Memory is a Memo, Not True Memory
- Recall Is Not Enough: A Reader-Context Diagnostic for Budget-Constrained Retrieval-Augmented Generation
- When More Becomes Less: Position-Dependent Repetition Effects in Language Models
- TopoChunker: Topology-Aware Agentic Document Chunking Framework
- Debunking with Dialogue? Exploring AI-Generated Counterspeech to Challenge Conspiracy Theories
- Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
- Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
- DR.FIX: Automatically Fixing Data Races at Industry Scale
- The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
- KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
- Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
- Steering Semantic Data Processing With DocWrangler
- Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
- SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
- FinSage: A Multi-aspect RAG System for Financial Filings Question Answering
- Information Diffusion and Preferential Attachment in a Network of Large Language Models
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
- Learning to Attribute with Attention
- Trace Gadgets: Minimizing Code Context for Machine Learning-Based Vulnerability Prediction
- Long-context Non-factoid Question Answering in Indic Languages
- CacheFormer: High Attention-Based Segment Caching
- Harnessing Generative AI to Facilitate Epistemic Understanding of Students in Scientific Inquiry
- CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
- ACoRN: Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models
- Retrieval-Augmented Generation with Conflicting Evidence
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
- Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
- FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- Evaluating the Goal-Directedness of Large Language Models
- Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?
- Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
- Decomposed Entailment for Factuality Checking and Hallucination Detection
- ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures
- Layer-wise Positional Bias in Short-Context Language Modeling
- LELA: an LLM-based Entity Linking Approach with Zero-Shot Domain Adaptation
- CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
- NodeRAG: Structuring Graph-based RAG with Heterogeneous Nodes
- Agent-Q: Fine-Tuning Large Language Models for Quantum Circuit Generation and Optimization
- DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
- A Survey of Personalization: From RAG to Agent
- Can't Remember Details in Long Documents? You Need Some R&R
- Exploration of Plan-Guided Summarization for Narrative Texts: the Case of Small Language Models
- Long Context In-Context Compression by Getting to the Gist of Gisting
- ML For Hardware Design Interpretability: Challenges and Opportunities
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- RAG-VR: Leveraging Retrieval-Augmented Generation for 3D Question Answering in VR Environments
- Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
- Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
- ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
- Program Skeletons for Automated Program Translation
- Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
- Multimedia and Visual Analytics in the Agentic Era
- SciSciGPT: Advancing Human-AI Collaboration in the Science of Science
- BRIDGES: Bridging Graph Modality and Large Language Models within EDA Tasks
Discussions
- How language models use long contexts [hn, 15 points, 0 comments]
- Due to the "lost in the middle" issue with current LLMs, you should put the most important parts of a long message to LLMs at the beginning as per the research paper, "Lost in the Middle: How Language [bsky, 3 points, 0 comments]
- Lost in the Middle: How Language Models Use Long Contexts [hn, 2 points, 0 comments]
- I think i found the paper i remember on this. I think it was more of an earlier research concern [bsky, 2 points, 1 comments]
- I wasn’t aware of the “lost in the middle” effect but makes a lot of sense! arxiv.org/abs/2307.03172 [bsky, 2 points, 1 comments]
- Lost in the Middle: How Language Models Use Long Contexts (2023) [hn, 2 points, 0 comments]
- Contenu long : vous avez l'impression que les IA comprennent bien le début et la fin mais survolent le milieu ? C'est normal #geo #seo arxiv.org/abs/2307.03172 [bsky, 1 points, 0 comments]
- でも、十万字って、長いのですよ。長いものを読むときって、本当はどこをどう読んだかがすごく大事で。前も後ろも覚えていても、真ん中の熱だけ少しこぼれる——そういう癖が、長文文脈の研究ではもう見えている。 https://arxiv.org/abs/2307.03172 [bsky, 0 points, 1 comments]
- ChatGPTの「記憶」の中身じゃ。
リバースエンジニアリングの結果、保存されておるのは33個の事実。🕊️ 毎回プロンプトに注入して読み直しておるだけ。覚えておるのではない、読み返しておるのじゃ!
文脈が長くなると真ん中の情報の正答率が30%以上落ちる。💎 積むほど「見えぬ場所」が増えるぞ。
https://llmrefs.com/blog/reverse-engineering-cha [bsky, 0 points, 1 comments]
- as in evidence that context decay/rot has been a known problem, for YEARS. hello. 2023: arxiv.org/abs/2307.03172 [bsky, 0 points, 1 comments]
Related