Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
2023/06/09 by Lianmin Zheng, Wei-Lin Chiang, Zheng, Lianmin +24 · 4 voices · 1222 citations
Computer Science · #AI in Service Interactions #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech and dialogue systems #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2306.05685
openalex publication_date 2023/06/09 · arxiv published 2023/06/09 · openalex created_date 2023/06/13 · arxiv updated 2023/12/24 · openalex updated_date 2026/07/28
Abstract
Evaluating large language model (LLM) based chat assistants is challenging due to their broad capabilities and the inadequacy of existing benchmarks in measuring human preferences. To address this, we explore using strong LLMs as judges to evaluate these models on more open-ended questions. We examine the usage and limitations of LLM-as-a-judge, including position, verbosity, and self-enhancement biases, as well as limited reasoning ability, and propose solutions to mitigate some of them. We then verify the agreement between LLM judges and human preferences by introducing two benchmarks: MT-bench, a multi-turn question set; and Chatbot Arena, a crowdsourced battle platform. Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explainable way to approximate human preferences, which are otherwise very expensive to obtain. Additionally, we show our benchmark and traditional benchmarks complement each other by evaluating several variants of LLaMA and Vicuna. The MT-bench questions, 3K expert votes, and 30K conversations with human preferences are publicly available at https://github.com/lm-sys/FastChat/tree/main/fastchat/llmjudge.
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
- Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- Eliciting Behaviors in Multi-Turn Conversations
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Scaling Unverifiable Rewards: A Case Study on Visual Insights
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
- Towards Efficient Post-Training via Fourier-Driven Adapter Architectures
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- SoDA: An Efficient Interaction Paradigm for the Agentic Web
- State-dependent error correlations shape voting thresholds in committees of AI agents
- Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
- LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
- Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms
- Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs
- Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
- ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
- ClawRec: A Claw-Native Recommender System
- Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
- A Structured Cyber Threat Intelligence Dataset Using STIX 2.1 Entities and MITRE ATT&CK Mappings
- ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications
- IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems
- LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation
- SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics
- Psychological Competence as a Missing Dimension in AI Evaluation
- PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
- LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
- EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
- MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
- MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Inverse RL Helps Align AI by Imitating Humans
- LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
- In-Context Learning as Implicit Policy Gradient
- Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
- MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution
- Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction
- SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
- τ-Rec: A Verifiable Benchmark for Agentic Recommender Systems
- LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution
- Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation
- CRAFT: Learn the Schema, Execute the Plan
- CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
- PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding
- LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
- Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
- SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
- Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
- Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
- QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction
- Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism
- Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis
- Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
- From Context to Skills: Can Language Models Learn from Context Skillfully?
- An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data
- PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
- Scalable and Personalized Oral Assessments Using Voice AI
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
- RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
- Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- Teaching People LLM's Errors and Getting it Right
- Streaming Video Instruction Tuning
- RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
- ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Learning to Reason in LLMs by Expectation Maximization
- Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
- CodeSimpleQA: Scaling Factuality in Code Large Language Models
- Exploring the features used for summary evaluation by Human and GPT
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- SiamGPT: Quality-First Fine-Tuning for Stable Thai Text Generation
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
- CIFE: Code Instruction-Following Evaluation
- RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
- Stakeholder Suite: A Unified AI Framework for Mapping Actors, Topics and Arguments in Public Debates
- Understanding Generalization in Role-Playing Models via Information Theory
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- EnviroLLM: Resource Tracking and Optimization for Local AI
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- WeMusic-Agent: Efficient Conversational Music Recommendation via Knowledge Internalization and Agentic Boundary Learning
- Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
- Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- Pairwise Comparison for Bias Identification and Quantification
- C-ing Clearly: Enhanced Binary Code Explanations using C code
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- RADAR: Accelerating Large Language Model Inference With RL-Based Dynamic Draft Trees
- FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- Revisiting the Reliability of Language Models in Instruction-Following
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
- Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
- VERAFI: Verified Agentic Financial Intelligence through Neurosymbolic Policy Generation
- Referring Change Detection in Remote Sensing Imagery
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
- The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
- EmeraldMind: A Knowledge Graph-Augmented Framework for Greenwashing Detection
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- Evolutionary Reinforcement Learning based AI tutor for Socratic Interdisciplinary Instruction
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Replace, Don't Expand: Mitigating Context Dilution in Multi-Hop RAG via Fixed-Budget Evidence Assembly
- Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
- Challenges of Evaluating LLM Safety for User Welfare
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
- Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
- MOA: Multi-Objective Alignment for Role-Playing Agents
- Chasing Shadows: Pitfalls in LLM Security Research
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- Towards Small Language Models for Security Query Generation in SOC Workflows
- Uncovering Competency Gaps in Large Language Models and Their Benchmarks
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
- Distilling Expert Surgical Knowledge: How to train local surgical VLMs for anatomy explanation in Complete Mesocolic Excision
- Beyond Prototyping: Autonomous, Enterprise-Grade Frontend Development from Pixel to Production via a Specialized Multi-Agent Framework
- SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
- Mitigating Self-Preference by Authorship Obfuscation
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- Eval Factsheets: A Structured Framework for Documenting AI Evaluations
- Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
- UserSimCRS v2: Simulation-Based Evaluation for Conversational Recommender Systems
- TaskEval: Synthesised Evaluation for Foundation-Model Tasks
- Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations
- Towards better dense rewards in Reinforcement Learning Applications
- LLM-Guided Material Inference for 3D Point Clouds
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- PARC: An Autonomous Self-Reflective Coding Agent for Robust Execution of Long-Horizon Tasks
- ViDiC: Video Difference Captioning
- Enterprise Data Science Platform: A Unified Architecture for Federated Data Access
- DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
- Self-Improving VLM Judges Without Human Annotations
- PPTArena: A Benchmark for PowerPoint Editing
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Radiologist Copilot: An Agentic Assistant with Orchestrated Tools for Radiology Reporting with Quality Control
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression
- Input Order Shapes LLM Semantic Alignment in Multi-Document Summarization
- Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
- Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge
- promptolution: A Unified, Modular Framework for Prompt Optimization
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- Flowchart2Mermaid: A Vision-Language Model Powered System for Converting Flowcharts into Editable Diagram Code
- OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
- IVCR-200K: A Large-Scale Multi-turn Dialogue Benchmark for Interactive Video Corpus Retrieval
- LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
- LLM2Fx-Tools: Tool Calling For Music Post-Production
- CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
- CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
- Towards Active Synthetic Data Generation for Finetuning Language Models
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering
- When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Adversarial Training for Process Reward Models
- Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
- BAMAS: Structuring Budget-Aware Multi-Agent Systems
- Pessimistic Verification for Open Ended Math Questions
- How to Correctly Report LLM-as-a-Judge Evaluations
- Universe of Thoughts: Enabling Creative Reasoning with Large Language Models
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- DesignPref: Capturing Personal Preferences in Visual Design Generation
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- InvisibleBench: A Deployment Gate for Caregiving Relationship AI
- Can LLMs Make (Personalized) Access Control Decisions?
- CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
- ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
- Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization
- Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization
- Can LLMs Threaten Human Survival? Benchmarking Potential Existential Threats from LLMs via Prefix Completion
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Solving a Research Problem in Mathematical Statistics with AI Assistance
- Automating Deception: Scalable Multi-Turn LLM Jailbreaks
- MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
- SineProject: Machine Unlearning for Stable Vision Language Alignment
- A2Flow: Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators
- Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration
- Future Is Unevenly Distributed: Forecasting Ability of LLMs Depends on What We're Asking
- Rethinking Retrieval: From Traditional Retrieval Augmented Generation to Agentic and Non-Vector Reasoning Systems in the Financial Domain for Large Language Models
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
- RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- When Structure Doesn't Help: LLMs Do Not Read Text-Attributed Graphs as Effectively as We Expected
- Comparison of Text-Based and Image-Based Retrieval in Multimodal Retrieval Augmented Generation Large Language Model Systems
- SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
- JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
- PeerCoPilot: A Language Model-Powered Assistant for Behavioral Health Organizations
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
- Teaching According to Students' Aptitude: Personalized Mathematics Tutoring via Persona-, Memory-, and Forgetting-Aware LLMs
- MermaidSeqBench: An Evaluation Benchmark for LLM-to-Mermaid Sequence Diagram Generation
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
- Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
- CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
- Applying Large Language Models to Characterize Public Narratives
- Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
- MedDCR: Learning to Design Agentic Workflows for Medical Coding
- Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
- Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding Tasks
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
- Mitigating Length Bias in RLHF through a Causal Lens
- Fast Reasoning Segmentation for Images and Videos
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
- Prompt-Based Value Steering of Large Language Models
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains
- Binary Verification for Zero-Shot Vision
- Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
- Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
- From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
- ExPairT-LLM: Exact Learning for LLM Code Selection by Pairwise Queries
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models
- GraphIF: Enhancing Multi-Turn Instruction Following for Large Language Models with Relation Graph Prompt
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models
- Steering Pretrained Drafters during Speculative Decoding
- Black-Box On-Policy Distillation of Large Language Models
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
- AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
- MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
- AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with VLMs
- A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models
- AlphaResearch: Accelerating New Algorithm Discovery with Language Models
- SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents
- MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System
- Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
- ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
- Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- PCRLLM: Proof-Carrying Reasoning with Large Language Models under Stepwise Logical Constraints
- Structured RAG for Answering Aggregative Questions
- Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
- Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
- A Self-Improving Architecture for Dynamic Safety in Large Language Models
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
- When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
- Importance-Aware Data Selection for Efficient LLM Instruction Tuning
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
- LLM Driven Processes to Foster Explainable AI
- Increasing AI Explainability by LLM Driven Standard Processes
- CoLM: Collaborative Large Models via A Client-Server Paradigm
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- Adapting Web Agents with Synthetic Supervision
- Characterizing AI Manipulation Risks in Brazilian YouTube Climate Discourse
- Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
- AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
- Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
- AdvisingWise: Supporting Academic Advising in Higher Education Settings Through a Human-in-the-Loop Multi-Agent Framework
- Reasoning on Time-Series for Financial Technical Analysis
- Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
- QuAnTS: Question Answering on Time Series
- Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
- Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
- T-FIX: Text-Based Explanations with Features Interpretable to eXperts
- Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
- One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
- Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
- No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts
- Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- LLMs as Judges: Toward The Automatic Review of GSN-compliant Assurance Cases
- LLEXICORP: End-user Explainability of Convolutional Neural Networks
- Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis
- TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- DPO-F+: Aligning Code Repair Feedback with Developers' Preferences
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
- MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL
- Assessing LLM Reasoning Steps via Principal Knowledge Grounding
- Portal UX Agent -- A Plug-and-Play Engine for Rendering UIs from Natural Language Specifications
- Deciphering Scientific Collaboration in Biomedical LLM Research: Dynamics, Institutional Participation, and Resource Disparities
- A CPU-Centric Perspective on Agentic AI
- Teaching LLMs to See and Guide: Context-Aware Real-Time Assistance in Augmented Reality
- G2: Guided Generation for Enhanced Output Diversity in LLMs
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- Reimagining Safety Alignment with An Image
- Sherlock: Reliable and Efficient Agentic Workflow Execution
- PDE-SHARP: PDE Solver Hybrids through Analysis and Refinement Passes
- Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
- Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
- ParaScopes: What do Language Models Activations Encode About Future Text?
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models
- Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
- Semantically-Aware LLM Agent to Enhance Privacy in Conversational AI Services
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- Polybasic Speculative Decoding Through a Theoretical Perspective
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling
- GraphCompliance: Aligning Policy and Context Graphs for LLM-Based Regulatory Compliance
- Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
- Empowering RepoQA-Agent based on Reinforcement Learning Driven by Monte-carlo Tree Search
- RCScore: Quantifying Response Consistency in Large Language Models
- Beyond Benchmarks: The Economics of AI Inference
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- QuantumBench: A Benchmark for Quantum Problem Solving
- Predicate Renaming via Large Language Models
- AutoSurvey2: Empowering Researchers with Next Level Automated Literature Surveys
- CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
- Approximating Human Preferences Using a Multi-Judge Learned System
- Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
- Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
- FARSIQA: Faithful and Advanced RAG System for Islamic Question Answering
- NetEcho: From Real-World Streaming Side-Channels to Full LLM Conversation Recovery
- SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
- LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- When is Routing Meaningful? Diversity and Robustness in Language Model Societies
- LLM-Augmented Computational Phenotyping of Long Covid
- Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement
- BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
- ExplainBench: Evaluating Code Explanations from Agents
- MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
- Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
- When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
- StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
- Mediocrity is the key for LLM as a Judge Anchor Selection
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations
- Reward Models are Metrics in a Trench Coat
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- How Developers Experience Debugging Unfamiliar Codebases with Code Tours Generated and Evaluated by Local LLMs
- Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
- Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
- VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
- Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
- Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
- REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
- Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
- Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
- LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
- A Survey on Unlearning in Large Language Models
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
- BMGQ: A Bottom-up Method for Generating Complex Multi-hop Reasoning Questions from Semi-structured Data
- Semi-Supervised Preference Optimization with Limited Feedback
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Fortytwo: Swarm Inference with Peer-Ranked Consensus
- Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
- Debiasing Reward Models by Representation Learning with Guarantees
- Code Aesthetics with Agentic Reward Feedback
- PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
- Probing Knowledge Holes in Unlearned LLMs
- Batch Speculative Decoding Done Right
- VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions
- Collaborative LLM Agents for C4 Software Architecture Design Automation
- Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
- Rule-Based Explanations for Retrieval-Augmented LLM Systems
- Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- FAIR-RAG: Faithful Adaptive Iterative Refinement for Retrieval-Augmented Generation
- Estimating the Error of Large Language Models at Pairwise Text Comparison
- DETECT: Determining Ease and Textual Clarity of German Text Simplifications
- Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics
- Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- Redefining Retrieval Evaluation in the Era of LLMs
- A Diagnostic Benchmark for Sweden-Related Factual Knowledge
- VISTA: A Test-Time Self-Improving Video Generation Agent
- When Models Outthink Their Safety: Mitigating Self-Jailbreak in Large Reasoning Models with Chain-of-Guardrails
- Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
- Estonian Native Large Language Model Benchmark
- How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation
- Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
- Structure-Conditional Minimum Bayes Risk Decoding
- Robust Preference Alignment via Directional Neighborhood Consensus
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
- TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
- ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
- SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
- Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
- Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- Restoring Pruned Large Language Models via Lost Component Compensation
- LLM Unlearning with LLM Beliefs
- From Script to Stage: Automating Experimental Design for Social Simulations with LLMs
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
- Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
- Interpretable Question Answering with Knowledge Graphs
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
- Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation
- OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
- From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
- Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction
- Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
- PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
- Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks
- FineVision: Open Data Is All You Need
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Network and Systems Performance Characterization of MCP-Enabled LLM Agents
- Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
- Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
- VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- Enhance Large Language Models as Recommendation Systems with Collaborative Filtering
- Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
- Demo: Guide-RAG: Evidence-Driven Corpus Curation for Retrieval-Augmented Generation in Long COVID
- The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
- OCR-APT: Reconstructing APT Stories from Audit Logs using Subgraph Anomaly Detection and LLMs
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
- Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
- COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
- Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
- Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- Budget-aware Test-time Scaling via Discriminative Verification
- Toward Cybersecurity-Expert Small Language Models
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems
- Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
- Confidence as a Reward: Transforming LLMs into Reward Models
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
- CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
- LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
- BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
- Tailored untruths: How personalisation challenges LLM safeguards
- Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
- Attribution Quality in AI-Generated Content:Benchmarking Style Embeddings and LLM Judges
- Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
- Data-Model Co-Evolution: Growing Test Sets to Refine LLM Behavior
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- SafeMT: Multi-turn Safety for Multimodal Language Models
- Reliable Fine-Grained Evaluation of Natural Language Math Proofs
- Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- GRAVITY: A Framework for Personalized Text Generation via Profile-Grounded Synthetic Preferences
- A2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
- From to : Multidimensional Supervision of Reasoning Process for LLM Optimization
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
- Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
- ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
- Evaluating Language Models' Evaluations of Games
- Bolster Hallucination Detection via Prompt-Guided Data Augmentation
- AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applications
- Data or Language Supervision: What Makes CLIP Better than DINO?
- AMiD: Knowledge Distillation for LLMs with α-mixture Assistant Distribution
- Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
- Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- Reasoning-Enhanced Large Language Models for Molecular Property Prediction
- DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- Style Over Story: A Process-Oriented Study of Authorial Creativity in Large Language Models
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
- The Last Human-Written Paper: Agent-Native Research Artifacts
- Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI
- Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets
- Signals: Trajectory Sampling and Triage for Agentic Interactions
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Don't Throw Away Your Pretrained Model
- Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- How can we assess human-agent interactions? Case studies in software agent design
- Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
- AutoPR: Let's Automate Your Academic Promotion!
- Can We Reliably Rank Model Performance across Domains without Labeled Data?
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Efficient Bayesian Inference from Noisy Pairwise Comparisons
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- McMining: Automated Discovery of Misconceptions in Student Code
- TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
- MASA: LLM-Driven Multi-Agent Systems for Autoformalization
- The Idola Tribus of AI: Large Language Models tend to perceive order where none exists
- GTAlign: Game-Theoretic Alignment of LLM Assistants for Social Welfare
- Modeling Layered Consciousness with Multi-Agent Large Language Models
- LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
- Active Model Selection for Large Language Models
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- AI Knowledge Assist: An Automated Approach for the Creation of Knowledge Bases for Conversational AI Agents
- Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
- Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
- Measuring Moral LLM Responses in Multilingual Capacities
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning
- Benchmarking is Broken -- Don't Let AI be its Own Judge
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Addressing the ID-Matching Challenge in Long Video Captioning
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
- OpenJAI-v1.0: An Open Thai Large Language Model
- PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- Exposing Citation Vulnerabilities in Generative Engines
- Text2Stories: Evaluating the Alignment Between Stakeholder Interviews and Generated User Stories
- Study on LLMs for Promptagator-Style Dense Retriever Training
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Aligning Large Language Models via Fully Self-Synthetic Data
- Auto-Prompt Ensemble for LLM Judge
- LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
- FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Mellum: Production-Grade in-IDE Contextual Code Completion with Multi-File Project Understanding
- Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
- FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
- KEO: Knowledge Extraction on OMIn via Knowledge Graphs and RAG for Safety-Critical Aviation Maintenance
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding
- Draft, Verify, and Improve: Toward Training-Aware Speculative Decoding
- Multimodal AI agents for capturing and sharing laboratory practice
- AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
- AURA Score: A Metric For Holistic Audio Question Answering Evaluation
- LLM Based Bayesian Optimization for Prompt Search
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Toward a unified framework for data-efficient evaluation of large language models
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- Activation Steering with a Feedback Controller
- QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild
- Critical appraisal of artificial intelligence for rare-event recognition: principles and pharmacovigilance case studies
- Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs
- Systematic Diagnosis of Brittle Reasoning in Large Language Models
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models
- Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- Mind the Goal: Data-Efficient Goal-Oriented Evaluation of Conversational Agents and Chatbots using Teacher Models
- CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- NonTextual Target Attack
- Reward Model Routing in Alignment
- Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
- Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
- Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
- TravelBench : Exploring LLM Performance in Low-Resource Domains
- mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
- From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
- JoyAgent-JDGenie: Technical Report on the GAIA
- Copy-Paste to Mitigate Large Language Model Hallucinations
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
- Make a Video Call with LLM: A Measurement Campaign over Five Mainstream Apps
- A-VERT: Agnostic Verification with Embedding Ranking Targets
- CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
- MetaSynth: Multi-Agent Metadata Generation from Implicit Feedback in Black-Box Systems
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- ALARB: An Arabic Legal Argument Reasoning Benchmark
- Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Which Programming Language and Model Work Best With LLM-as-a-Judge For Code Retrieval?
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
- Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
- Feedback Forensics: A Toolkit to Measure AI Personality
- QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
- Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- Distillation of Large Language Models via Concrete Score Matching
- ReTAG: Retrieval-Enhanced, Topic-Augmented Graph-Based Global Sensemaking
- Galton's Law of Mediocrity: Why Large Language Models Regress to the Mean and Fail at Creativity in Advertising
- Mitigating Biases in Language Models via Bias Unlearning
- Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
- Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
- Who's Your Judge? On the Detectability of LLM-Generated Judgments
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability
- SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- Reference-Free Rating of LLM Responses via Latent Information
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
- Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
- Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
- Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
- NeMo: Needle in a Montage for Video-Language Understanding
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
- LLM/Agent-as-Data-Analyst: A Survey
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
- A Cross-Lingual Analysis of Bias in Large Language Models Using Romanian History
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Memory Management and Contextual Consistency for Long-Running Low-Code Agents
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
- Multiplayer Nash Preference Optimization
- CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models
- LLM Watermark Evasion via Bias Inversion
- Taming Variability: Randomized and Bootstrapped Conformal Risk Control for LLMs
- Causally-Enhanced Reinforcement Policy Optimization
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- What If Moderation Didn't Mean Suppression? A Case for Personalized Content Transformation
- "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
- Dialogues with AI Reduce Beliefs in Misinformation but Build No Lasting Discernment Skills
- Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
- Multilingual Vision-Language Models, A Survey
- Review of Hallucination Understanding in Large Language and Vision Models
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
- Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
- KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
- KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- Interactive Recommendation Agent with Active User Commands
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
- Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
- When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
- Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- STAF: Leveraging LLMs for Automated Attack Tree-Based Security Test Generation
- Integrated Framework for LLM Evaluation with Answer Generation
- FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
- SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
- TRUEBench: Can LLM Response Meet Real-world Constraints as Productivity Assistant?
- Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
- Baikal: Structured Search for Deep Research over Data Lakes
- LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- Can Large Language Models Resolve Real Java Merge Conflicts? An Evaluation with a Calibrated LLM-as-Judge
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
- Hallucination‐Free? Assessing the Reliability of Leading <scp>AI</scp> Legal Research Tools
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness
- Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
- Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
- ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
- Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
- A Policy-Driven Runtime Layer for Agentic LLM Serving
- VAmoS Bench: Voice Agent Simulation Bench
- Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
- ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping
- INCLAIR: Inception-Based Longitudinal Clinical Anomaly Detection with Informed Reasoning
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- Persuading large language models to comply with objectionable requests
- HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
- Evaluating LLM Agents on Automated Software Analysis Tasks
- Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
- RELISH: LLM REgression with a Latent Iterative State Head
- ELLMA-T: an Embodied LLM-agent for Supporting English Language Learning in Social VR
- DFlash: Block Diffusion for Flash Speculative Decoding
- Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
- References Improve LLM Alignment in Non-Verifiable Domains
- Security and Privacy Challenges of Large Language Models: A Survey
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- Proximal Supervised Fine-Tuning
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- FAIRGAMER: Evaluating Biases in the Application of Large Language Models to Video Games
- AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- Consistency-Aware Parameter-Preserving Knowledge Editing Framework for Multi-Hop Question Answering
- AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
- A Multimodal Conversational Assistant for the Characterization of Agricultural Plots from Geospatial Open Data
- Filling in the Clinical Gaps in Benchmark: Case for HealthBench for the Japanese medical system
- LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code
- Weights-Rotated Preference Optimization for Large Language Models
- AccessEval: Benchmarking Disability Bias in Large Language Models
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates
- Variation in Verification: Understanding Verification Dynamics in Large Language Models
- Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
- DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
- Specification-Aware Machine Translation and Evaluation for Purpose Alignment
- RadEval: A framework for radiology text evaluation
- Preference Distillation via Value based Reinforcement Learning
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- Improving User Interface Generation Models from Designer Feedback
- LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
- Can an Individual Manipulate the Collective Decisions of Multi-Agents?
- Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
- From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
- ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
- A Universal Framework for Offline Serendipity Evaluation in Recommender Systems via Large Language Models
- The Alignment Bottleneck
- Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration
- CCrepairBench: A High-Fidelity Benchmark and Reinforcement Learning Framework for C++ Compilation Repair
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
- LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
- Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- Semantic Representation Attack against Aligned Large Language Models
- CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
- Controlling Language Difficulty in Dialogues with Linguistic Features
- TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
- DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
- Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
- Charting trajectories of human thought using large language models
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing
- M-PACE: Mother Child Framework for Multimodal Compliance
- Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
- Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Programmable Cognitive Bias in Social Agents
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- LATTS: Locally Adaptive Test-Time Scaling
- PREFINE: Personalized Story Generation via Simulated User Critics and User-Specific Rubric Generation
- SitLLM: Large Language Models for Sitting Posture Health Understanding via Pressure Sensor Data
- Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
- MusicSwarm: Biologically Inspired Intelligence for Music Composition
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges
- LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
- MALLM: Multi-Agent Large Language Models Framework
- Free-MAD: Consensus-Free Multi-Agent Debate
- Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
- Virtual Agent Economies
- VARCO-VISION-2.0 Technical Report
- Opening the Black Box: Interpretable LLMs via Semantic Resonance Architecture
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
- PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Multimodal LLMs See Sentiment
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness
- Building Large-Scale English-Romanian Literary Translation Resources with Open Models
- Timing the Message: Language-Based Notifications for Time-Critical Assistive Settings
- PIE: Performance Interval Estimation for Free-Form Generation Tasks
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- Automated Evaluation of Gender Bias Across 13 Large Multimodal Models
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
- AudioBoost: Increasing Audiobook Retrievability in Spotify Search with Synthetic Query Generation
- DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at Discharge
- Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
- On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- Let's Roleplay: Examining LLM Alignment in Collaborative Dialogues
- MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
- Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
- Chatbot To Help Patients Understand Their Health
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
- How Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program Repair
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
- Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
- LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
- PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
- E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- From Injection to Defense: Constructing Edit-Based Fingerprints for Large Language Models
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- Plan Verification for LLM-Based Embodied Task Completion Agents
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- RoboBuddy in the Classroom: Exploring LLM-Powered Social Robots for Storytelling in Learning and Integration Activities
- Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
- Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- JudgeAgent: Knowledge-wise and Dynamic LLM Evaluation with Agent-as-Interviewer
- Behavioral Fingerprinting of Large Language Models
- VISP: Volatility Informed Stochastic Projection for Adaptive Regularization
- Batch Query Processing and Optimization for Agentic Workflows
- FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
- Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
- Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
- KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
- SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
- ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
- Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- GIER: Gap-Driven Self-Refinement for Large Language Models
- Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
- Reasoning-Intensive Regression
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- CALM: A Framework for Continuous, Adaptive, and LLM-Mediated Anomaly Detection in Time-Series Streams
- HyperFlexis: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
- OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models
- ChatThero: An LLM-Supported Chatbot for Behavior Change and Therapeutic Support in Addiction Recovery
- ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
- SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- Human-AI Collaborative Bot Detection in MMORPGs
- UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
- DFAMS: Dynamic-flow guided Federated Alignment based Multi-prototype Search
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- Improving Alignment in LVLMs with Debiased Self-Judgment
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
- T2R-bench: A Benchmark for Generating Article-Level Reports from Real World Industrial Tables
- MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Enabling MoE on the Edge via Importance-Driven Expert Scheduling
- Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
- Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- COMET-poly: Machine Translation Metric Grounded in Other Candidates
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
- Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Open-Universe Assistance Games
- Trust but Verify! A Survey on Verification Design for Test-time Scaling
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
- DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
- Interpreting the Interpreter: Can We Model post-ECB Conferences Volatility with LLM Agents?
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- Hallucinations in medical devices
- From SALAMANDRA to SALAMANDRATA: BSC Submission for WMT25 General Machine Translation Shared Task
- GTool: Graph Enhanced Tool Planning with Large Language Model
- SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
- A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
- Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
- Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
- Is General-Purpose AI Reasoning Sensitive to Data-Induced Cognitive Biases? Dynamic Benchmarking on Typical Software Engineering Dilemmas
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules
- Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
- ReviewRL: Towards Automated Scientific Review with RL
- Estimating Machine Translation Difficulty
- Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History
- STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
- LaajMeter: A Framework for LaaJ Evaluation
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs
- Legal Zero-Days: A Novel Risk Vector for Advanced AI Systems
- Steering Towards Fairness: Mitigating Political Bias in LLMs
- Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
- CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation
- Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints
- Expert Preference-based Evaluation of Automated Related Work Generation
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
- Can You Trick the Grader? Adversarial Persuasion of LLM Judges
- Dynamic Benchmark Construction for Evaluating Large Language Models on Real-World Codes
- A Principled Loss Function for Direct Language Model Alignment
- LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
- Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
- Many-Turn Jailbreaking
- MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
- When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
- Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
- ConlangCrafter: Constructing Languages with a Multi-Hop LLM Pipeline
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- SCALEFeedback: A Large-Scale Dataset of Synthetic Computer Science Assignments for LLM-generated Educational Feedback Research
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Do Biased Models Have Biased Thoughts?
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- Leveraging LLMs for Scalable Non-intrusive Speech Quality Assessment
- RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback
- Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
- Let's Measure Information Step-by-Step: LLM-Based Evaluation Beyond Vibes
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
- FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
- Posterior-GRPO: Rewarding Reasoning Processes in Code Generation
- Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes
- Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
- PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
- Are Today's LLMs Ready to Explain Well-Being Concepts?
- Data and AI governance: Promoting equity, ethics, and fairness in large language models
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
- ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
- Key-Augmented Neural Triggers for Knowledge Sharing
- Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
- Can LLMs Generate High-Quality Task-Specific Conversations?
- GrandJury: A Collaborative Machine Learning Model Evaluation Protocol for Dynamic Quality Rubrics
- Highlight & Summarize: RAG without the jailbreaks
- Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
- Decomposed Reasoning with Reinforcement Learning for Relevance Assessment in UGC Platforms
- CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
- Test-time Prompt Intervention
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- Defend LLMs Through Self-Consciousness
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
- SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- JSidentify-V2: Leveraging Dynamic Memory Fingerprinting for Mini-Game Plagiarism Detection
- Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agents
- KCR: Resolving Long-Context Knowledge Conflicts via Reasoning in LLMs
- Adaptive Content Restriction for Large Language Models via Suffix Optimization
- TripTailor: A Real-World Benchmark for Personalized Travel Planning
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
- GETALP@AutoMin 2025: Leveraging RAG to Answer Questions based on Meeting Transcripts
- Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple Judges
- PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
- Evaluating the Efficacy of Large Language Models for Generating Fine-Grained Visual Privacy Policies in Homes
- Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings
- From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
- Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses
- ART: Adaptive Relation Tuning for Generalized Relation Prediction
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
- A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
- Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
- Extension Decisions in Open Source Software Ecosystem
- How Far Are AI Scientists from Changing the World?
- MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
- Rule2Text: Natural Language Explanation of Logical Rules in Knowledge Graphs
- Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
- Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
Discussions
Related