Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1347 citations
#cs.CL #cs.AI
paper · pdf · doi:10.48550/arxiv.2507.06261
Abstract
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.
Cited by
- Benchmarking Text-to-SQL under Role-Based Access Control
- MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery
- SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
- IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
- GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG
- ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
- Local Brushstroke Quality Assessment via Vision-Language Feedback
- DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
- RS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image Editing
- TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
- Experience Augmented Policy Optimization for LLM Reasoning
- WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
- TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code Development
- GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models
- M-RAG: Semantic Key-Value Indexing for Retrieval-Augmented Generation
- D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
- Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
- Why Do Vision Language Models Struggle To Recognize Human Emotions?
- Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
- GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
- Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
- Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
- Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
- CGCE: Classifier-Guided Concept Erasure in Generative Models
- Meta-Learning Preferences for Multilingual LLM Alignment
- ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
- Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
- FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
- Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering
- Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
- Perturbation is All You Need for Extrapolating Language Models
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
- When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
- H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
- RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
- From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
- LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
- B-repLer: Language-guided Editing of CAD Models
- Is "Knowing It's Malicious Enough?" Evaluating LLMs for Fine-Grained Malware Behavior Auditing
- CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
- Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
- Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents
- AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
- CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
- RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
- Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
- LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
- Knowledge-Centric Agents for Workflow Generation in ComfyUI
- KD-Judge: A Knowledge-Driven Automated Judge Framework for Functional Fitness Movements on Edge Devices
- Native Active Perception as Reasoning for Omni-Modal Understanding
- SceneBind: Binding What and Where Across Vision, Audio and Language
- VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
- VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
- RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
- MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
- ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
- Fully Automated End-to-End Adversary Emulation from MITRE ATT&CK Based Cyber Threat Intelligence Using LLMs
- Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
- Segmenting Human-LLM Co-authored Text via Change Point Detection
- TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
- Gemma 4 Technical Report
- AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
- Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
- Neutrality Bites: Gender Representation in AI-Generated Animal Stories
- Lost in Context: Addressing Context Anxiety in Large Language Models
- CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation
- The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
- LLMs Corrupt Your Documents When You Delegate
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
- Can LLMs Beat Classical Hyperparameter Optimization Algorithms? A Study on autoresearch
- Greater accessibility can amplify discrimination in generative AI
- MIRAGE: The Illusion of Visual Understanding
- Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
- Position: Modular Memory is the Key to Continual Learning Agents
- Discovering Multiagent Learning Algorithms with Large Language Models
- Memory Caching: RNNs with Growing Memory
- Linear representations in language models can change dramatically over a conversation
- Replicating Human Motivated Reasoning Studies with LLMs
- Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
- Measuring the State of Open Science in Transportation Using Large Language Models
- Reasoning Models Ace the CFA Exams
- TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
- TRINITY: An Evolved LLM Coordinator
- Learning to Orchestrate Agents in Natural Language with the Conductor
- Echoing: Identity Failures when LLM Agents Talk to Each Other
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
- Accumulating Context Changes the Beliefs of Language Models
- Fit for Purpose? Deepfake Detection in the Real World
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- EditLens: Quantifying the Extent of AI Editing in Text
- Video models are zero-shot learners and reasoners
- The Illusion of Readiness in Health AI
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- On the Theoretical Limitations of Embedding-Based Retrieval
- Privileged Self-Access Matters for Introspection in AI
- Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
- Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- UCAgents: Unidirectional Convergence for Visual Evidence Anchored Multi-Agent Medical Decision-Making
- The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
- From Generated Human Videos to Physically Plausible Robot Trajectories
- Qwen-Music Technical Report
- Active Perception Agent for Omnimodal Audio-Video Understanding
- AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
- ProGuard: Towards Proactive Multimodal Safeguard
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
- MetFuse: Figurative Fusion between Metonymy and Metaphor
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- PaperBanana: Automating Academic Illustration for AI Scientists
- Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
- Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
- HiFi-RAG: Hierarchical Content Filtering and Two-Pass Generation for Open-Domain RAG
- AnalogSAGE: Self-evolving Analog Design Multi-Agents with Stratified Memory and Grounded Experience
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
- LiveProteinBench: A Contamination-Free Benchmark for Assessing Models' Specialized Capabilities in Protein Science
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation
- Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
- Offline-Online Curriculum RL for Multimodal Reasoning
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications
- SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
- CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
- ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
- Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management
- Inverse RL Helps Align AI by Imitating Humans
- In-Context Learning as Implicit Policy Gradient
- DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video
- Layering Virtual Try-On
- Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records
- ID-V2V: Identity-Preserving Video Restylization
- Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning
- LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution
- VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
- Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
- Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
- What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization
- Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
- Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
- PRIMA: Pre-Training with Risk-Integrated Image--Metadata Alignment for Medical Diagnosis with LLM-Based Feature Aggregation
- SVBench: Evaluation of Video Generation Models on Social Reasoning
- LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- TrafficSimAgent: A Hierarchical Agent Framework for Autonomous Traffic Simulation with MCP Control
- Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation
- ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
- TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
- H2em: Learning Hierarchical Hyperbolic Embeddings for Compositional Zero-Shot Learning
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
- Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
- From Retrieval to Reasoning: A Framework for Cyber Threat Intelligence NER with Explicit and Adaptive Instructions
- Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
- DeliveryBench: Can Agents Earn Profit in Real World?
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement Learning
- Does It Tie Out? Towards Autonomous Legal Agents in Venture Capital
- OpenView: Empowering MLLMs with Out-of-view VQA
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- Name That Part: 3D Part Segmentation and Naming
- When Reasoning Meets Its Laws
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Xiaomi MiMo-VL-Miloco Technical Report
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
- A systematic assessment of Large Language Models for constructing two-level fractional factorial designs
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- Scaling Laws for Energy Efficiency of Local LLMs
- VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
- Large Video Planner Enables Generalizable Robot Control
- You Never Know a Person, You Only Know Their Defenses: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations
- Corrective Diffusion Language Models
- GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
- Step-GUI Technical Report
- FAME: Fictional Actors for Multilingual Erasure
- CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- Will AI Trade? A Computational Inversion of the No-Trade Theorem
- DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
- MedNuggetizer: Confidence-Based Information Nugget Extraction from Medical Documents
- V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- Visual-textual Dermatoglyphic Animal Biometrics: A First Case Study on Panthera tigris
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
- ART: Articulated Reconstruction Transformer
- Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- Semantic search for 100M+ galaxy images using AI-generated captions
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis
- Grammar Search for Multi-Agent Systems
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- MMhops-R1: Multimodal Multi-hop Reasoning
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- Towards Interactive Intelligence for Digital Humans
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
- LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
- MineTheGap: Automatic Mining of Biases in Text-to-Image Models
- State over Tokens: Characterizing the Role of Reasoning Tokens
- Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- AutoMV: An Automatic Multi-Agent System for Music Video Generation
- Training Versatile Coding Agents in Synthetic Environments
- BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- Using GUI Agent for Electronic Design Automation
- Reconstruction as a Bridge for Event-Based Visual Question Answering
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
- FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
- Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- LLMs Can Assist with Proposal Selection at Large User Facilities
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
- Agile Deliberation: Concept Deliberation for Subjective Visual Classification
- Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
- ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation
- Evaluating Gemini Robotics Policies in a Veo World Simulator
- Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
- Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
- ChronusOmni: Improving Time Awareness of Omni Large Language Models
- CNFinBench: A Benchmark for Safety and Compliance of Large Language Models in Finance
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- SoMe: A Realistic Benchmark for LLM-based Social Media Agents
- From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model
- Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
- Relational Visual Similarity
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
- In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection
- Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
- NeSTR: A Neuro-Symbolic Abductive Framework for Temporal Reasoning in Large Language Models
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- Agency at the Interface: Distinguishing Teleological from Structural Self-Organization via Internal Coarse-Graining and Downward Causation
- MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
- Ideal Attribution and Faithful Watermarks for Language Models
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Small Language Models Can Use Nuanced Reasoning For Health Science Research Classification: A Microbial-Oncogenesis Case Study
- Enhanced Multimodal Video Retrieval System: Integrating Query Expansion and Cross-modal Temporal Event Retrieval
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- Distilling Expert Surgical Knowledge: How to train local surgical VLMs for anatomy explanation in Complete Mesocolic Excision
- Training Multi-Image Vision Agents via End2End Reinforcement Learning
- Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
- The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
- Are Your Agents Upward Deceivers?
- Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
- Qwen3.5-Omni Technical Report
- Verbalizing LLMs' assumptions to explain and control sycophancy
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
- Thinking with Programming Vision: Towards a Unified View for Thinking with Images
- Refaçade: Editing Object with Given Reference Texture
- Not All Birds Look The Same: Identity-Preserving Generation For Birds
- Open-Ended Goal Inference through Actions and Language for Human-Robot Collaboration
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
- 6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
- Unique Lives, Shared World: Learning from Single-Life Videos
- PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design
- Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- Automatic Attack Discovery for Few-Shot Class-Incremental Learning via Large Language Models
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
- OneThinker: All-in-one Reasoning Model for Image and Video
- PPTArena: A Benchmark for PowerPoint Editing
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- LORE: A Large Generative Model for Search Relevance
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
- Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge
- Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- VACoT: Rethinking Visual Data Augmentation with VLMs
- OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- Rectifying LLM Thought from Lens of Optimization
- OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
- Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability
- CauSight: Learning to Supersense for Visual Causal Discovery
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing
- SynthStrategy: Extracting and Formalizing Latent Strategic Insights from LLMs in Organic Chemistry
- ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
- FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting
- CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
- Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
- SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- LLM2Fx-Tools: Tool Calling For Music Post-Production
- CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
- ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering
- Design and Evaluation of a Multi-Agent Perception System for Autonomous Flying Networks
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- GreenPlanner: Practical Floorplan Layout Generation via an Energy-Aware and Function-Feasible Generative Framework
- When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
- IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
- VCWorld: A Biological World Model for Virtual Cell Simulation
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- CodeFlowLM: Incremental Just-In-Time Defect Prediction with Pretrained Language Models and Exploratory Insights into Defect Localization
- Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
- ThetaEvolve: Test-time Learning on Open Problems
- HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
- MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
- Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
- Geometrically-Constrained Agent for Spatial Reasoning
- DisCEdge: Distributed Context Management for Large Language Models at the Edge
- RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
- Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- Seeing without Pixels: Perception from Camera Trajectories
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Co-Evolving Agents: Learning from Failures as Hard Negatives
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- On the Limits of Innate Planning in Large Language Models
- Qwen3-VL Technical Report
- Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- Layer-Aware Video Composition via Split-then-Merge
- Vision-Language Memory for Spatial Reasoning
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
- Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Block Cascading: Training Free Acceleration of Block-Causal Video Models
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
- CREward: A Type-Specific Creativity Reward Model
- OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
- Low-Resolution Editing is All You Need for High-Resolution Editing
- Simulated Self-Assessment in Large Language Models: A Psychometric Approach to AI Self-Efficacy
- MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorization
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization
- HunyuanOCR Technical Report
- PRInTS: Reward Modeling for Long-Horizon Information Seeking
- AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
- Vidi2: Large Multimodal Models for Video Understanding and Creation
- Addressing Situated Teaching Needs: A Multi-Agent Framework for Automated Slide Adaptation
- Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
- RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry
- Plan-X: Instruct Video Generation via Semantic Planning
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation
- Native 3D Editing with Full Attention
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
- Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
- OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- Evaluating Adversarial Vulnerabilities in Modern Large Language Models
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
- Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated Narrations
- FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
- Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints
- ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025
- OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
- Step-Audio-R1 Technical Report
- Think Visually, Reason Textually: Vision-Language Synergy in ARC
- First Frame Is the Place to Go for Video Content Customization
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
- ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
- Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- Insert In Style: A Zero-Shot Generative Framework for Harmonious Cross-Domain Object Composition
- UniSER: A Foundation Model for Unified Soft Effects Removal
- Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
- GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- FxSearcher: gradient-free text-driven audio transformation
- Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
- P1: Mastering Physics Olympiads with Reinforcement Learning
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Video Spatial Reasoning with Object-Centric 3D Rollout
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding Tasks
- Direct Visual Grounding by Directing Attention of Visual Tokens
- BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
- Explainable AI-Generated Image Detection RewardBench
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
- RTMol: Rethinking Molecule-text Alignment in a Round-trip View
- Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- Scalable Policy Evaluation with Video World Models
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- STaR: Towards Cognitive Table Reasoning via Slow-Thinking Large Language Models
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- Documenting SME Processes with Conversational AI: From Tacit Knowledge to BPMN
- Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
- SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
- Mastering Olympiad-Level Physics with Artificial Intelligence
- Black-Box On-Policy Distillation of Large Language Models
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
- CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
- The 2025 Planning Performance of Frontier Large Language Models
- ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- AlphaCast: A Human Wisdom-LLM Intelligence Co-Reasoning Framework for Interactive Time Series Forecasting
- mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models
- AlphaResearch: Accelerating New Algorithm Discovery with Language Models
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
- UI2CodeN: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation
- Estranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
- State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
- Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- Streaming Tensor Program: A streaming abstraction for dynamic parallelism
- UniREditBench: A Unified Reasoning-based Image Editing Benchmark
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- TreeWriter: AI-Assisted Hierarchical Planning and Writing for Long-Form Documents
- FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
- GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning
- Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
- Using Language Models as Closed-Loop High-Level Planners for Robotics Applications: A Brief Overview and Benchmarks
- SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
- MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
- JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
- NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation
- In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- Retrieval Quality at Context Limit
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- Culture in Action: Evaluating Text-to-Image Models through Social Activities
- The Peril of Preference: Why GRPO fails on Ordinal Rewards
- Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
- iFlyBot-VLM Technical Report
- Scientific judgment drifts over time in AI ideation
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
- Motif 2 12.7B technical report
- Visual Spatial Tuning
- Quantifying the Climate Risk of Generative AI: Region-Aware Carbon Accounting with G-TRACE and the AI Sustainability Pyramid
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Modeling Clinical Uncertainty in Radiology Reports: from Explicit Uncertainty Markers to Implicit Reasoning Pathways
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
- T-FIX: Text-Based Explanations with Features Interpretable to eXperts
- Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
- To See or To Read: User Behavior Reasoning in Multimodal LLMs
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels
- Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- InsurAgent: A Large Language Model-Empowered Agent for Simulating Individual Behavior in Purchasing Flood Insurance
- Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images
- Driving scenario generation and evaluation using a structured layer representation and foundational models
- FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
- VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
- MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
- Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
- LongCat-Flash-Omni Technical Report
- SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
- Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
- NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
- Scaling Image Geo-Localization to Continent Level
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- Evaluating Perspectival Biases in Cross-Modal Retrieval
- Emu3.5: Native Multimodal Models are World Learners
- InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
- Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration
- ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
- CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
- WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
- Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
- VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
- Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
- Reward Models are Metrics in a Trench Coat
- CoDA: Agentic Systems for Collaborative Data Visualization
- COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
- AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World
- Autoregressive Boltzmann Generators
- VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
- Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
- A Language for Describing Agentic LLM Contexts
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
- Evaluation of Agents under Simulated AI Marketplace Dynamics
- SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
- MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
- SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations
- Prior Knowledge Makes It Possible: From Sublinear Graph Algorithms to LLM Test-Time Methods
- Atom-anchored LLMs speak Chemistry: A Retrosynthesis Demonstration
- Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
- Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech
- Magellan: Autonomous Discovery of Novel Compiler Optimization Heuristics with AlphaEvolve
- TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
- KnowCoder-A1: Incentivizing Agentic Reasoning Capability with Outcome Supervision for KBQA
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
- Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
- OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
- TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- ChessQA: Evaluating Large Language Models for Chess Understanding
- CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
- Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
- ATA: A Neuro-Symbolic Approach to Implement Autonomous and Trustworthy Agents
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- ReCode: Unify Plan and Action for Universal Granularity Control
- ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
- SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
- Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
- Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
- VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- Evaluating Multimodal Large Language Models on Core Music Perception Tasks
- PerCoR: Evaluating Commonsense Reasoning in Persian via Multiple-Choice Sentence Completion
- UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- SteerX: Disentangled Steering for LLM Personalization
- PACR: Progressively Ascending Confidence Reward for LLM Reasoning
- LightAgent: Mobile Agentic Foundation Models
- TripTide: A Benchmark for Adaptive Travel Planning under Disruptions
- L2M3OF: A Large Language Multimodal Model for Metal-Organic Frameworks
- Video-As-Prompt: Unified Semantic Control for Video Generation
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Generalizable Reasoning through Compositional Energy Minimization
- ComProScanner: A multi-agent based framework for composition-property structured data extraction from scientific literature
- A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- Rethinking Cross-lingual Gaps from a Statistical Viewpoint
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
- Black Box Absorption: LLMs Undermining Innovative Ideas
- Code-enabled language models can outperform reasoning models on diverse tasks
- LayerComposer: Multi-Human Personalized Generation via Layered Canvas
- Analyticup E-commerce Product Search Competition Technical Report from Team TredenceAICOE
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Exploring Conditions for Diffusion models in Robotic Control
- Data-Centric Lessons To Improve Speech-Language Pretraining
- TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- LLM Unlearning with LLM Beliefs
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
- News-Aware Direct Reinforcement Trading for Financial Markets
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Extracting alignment data in open models
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
- Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Glyph: Scaling Context Windows via Visual-Text Compression
- LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
- BenCao: An Instruction-Tuned Large Language Model for Traditional Chinese Medicine
- DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
- LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
- Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Physics-Informed Large Language Models for HVAC Anomaly Detection with Autonomous Rule Generation
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- LexChain: Modeling Legal Reasoning Chains for Chinese Tort Case Analysis
- Real-Time World Crafting: Generating Structured Game Behaviors from Natural Language with Large Language Models
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
- PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
- DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
- Composition-Grounded Instruction Synthesis for Visual Reasoning
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- Coder as Editor: Code-driven Interpretable Molecular Optimization
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- ZeroFalse: Improving Precision in Static Analysis with LLMs
- UniCode: A Framework for Generating High Quality Competitive Coding Problems
- PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
- Qwen3Guard Technical Report
- Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- CodeEvolve: An open source evolutionary coding agent for algorithm discovery and optimization
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- BioMedSearch: A Multi-Source Biomedical Retrieval Framework Based on LLMs
- StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- The Mechanistic Emergence of Symbol Grounding in Language Models
- Generative Universal Verifier as Multimodal Meta-Reasoner
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- Adaptive Reasoning Executor: A Collaborative Agent System for Efficient Reasoning
- BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Detect Anything via Next Point Prediction
- Data-Model Co-Evolution: Growing Test Sets to Refine LLM Behavior
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- Diff-XYZ: A Benchmark for Evaluating Diff Understanding
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- HoneyBee: Data Recipes for Vision-Language Reasoners
- Scaling Language-Centric Omnimodal Representation Learning
- Inferring Dynamic Physical Properties from Video Foundation Models
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
- ParaCook: On Time-Efficient Planning for Multi-Agent Systems
- GlobalizeEd: A Multimodal Translation System that Preserves Speaker Identity in Academic Lectures
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- What Generative Search Engines Like and How to Optimize Web Content Cooperatively
- Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems
- Neural Weight Compression for Language Models
- Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- A Survey on Agentic Multimodal Large Language Models
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
- BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- AwareCompiler: Agentic Context-Aware Compiler Optimization via a Synergistic Knowledge-Data Driven Framework
- SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn AGent Evaluation
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
- AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
- EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
- Agro-Consensus: Semantic Self-Consistency in Vision-Language Models for Crop Disease Management in Developing Countries
- Don't Just Fine-tune the Agent, Tune the Environment
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
- Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- Deliberative Dynamics and Value Alignment in LLM Debates
- An agentic artificially intelligent X-ray scientist
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
- AutoPR: Let's Automate Your Academic Promotion!
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- Towards a Taxonomy of Sustainability Requirements for Software Design
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- SEER: Sustainability Enhanced Engineering of Software Requirements
- DARO: Difficulty-Aware Reweighting Policy Optimization
- ENLighten: Lighten the Transformer, Enable Efficient Optical Acceleration
- Symskill: Symbol and Skill Co-Invention for Data-Efficient and Real-Time Long-Horizon Manipulation
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
- Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
- CREST-Search: Comprehensive Red-teaming for Evaluating Safety Threats in Large Language Models Powered by Web Search
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
- Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
- MeSH: Memory-as-State-Highways for Recursive Transformers
- LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
- DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- Neologism Learning for Controllability and Self-Verbalization
- Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
- LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
- VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
- Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
- LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
- Efficient numeracy in language models through single-token number embeddings
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- Online Rubrics Elicitation from Pairwise Comparisons
- COMPASS: Benchmarking Constrained Optimization in LLM Agents
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
- Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
- TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
- Active Semantic Perception
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Finish First, Perfect Later: Test-Time Token-Level Cross-Validation for Diffusion Large Language Models
- AI-to-AI Feedback: Amplified Intelligence — Prior Art and Governance Implications for Multi-Model Advisory Architectures
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
- Detecting Distillation Data from Reasoning Models
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Robustness assessment of large audio language models in multiple-choice evaluation
- Conditional Representation Learning for Customized Tasks
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
- TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- A Lightweight Large Language Model-Based Multi-Agent System for 2D Frame Structural Analysis
- SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Learning from All: Concept Alignment for Autonomous Distillation from Multiple Drifting MLLMs
- Toward a unified framework for data-efficient evaluation of large language models
- Quantitative Certification of Agentic Tool Selection
- CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
- When and Where do Events Switch in Multi-Event Video Generation?
- NonTextual Target Attack
- AudioToolAgent: An Agentic Framework for Audio-Language Models
- A Granular Study of Safety Pretraining under Model Abliteration
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
- Hearing the Order: Investigating Selection Bias in Large Audio-Language Models
- When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
- Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
- AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance
- ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- Automated Structured Radiology Report Generation with Rich Clinical Context
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- Cutting the Skip: Training Residual-Free Transformers
- DiSC-AMC: Token- and Parameter-Efficient Discretized Statistics In-Context Automatic Modulation Classification
- Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
- Generating Difficult-to-Translate Texts
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Towards Verified Code Reasoning by LLMs
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
- SCUBA: Salesforce Computer Use Benchmark
- Game-Time: Evaluating Temporal Dynamics in Spoken Language Models
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- Text-to-Scene with Large Reasoning Models
- Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
- Better Privilege Separation for Agents by Restricting Data Types
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
- Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Distillation of Large Language Models via Concrete Score Matching
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- DescribeEarth: Describe Anything for Remote Sensing Images
- Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- InfoAgent: Advancing Autonomous Information-Seeking Agents
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
- Agentic Exploration of Physics Models
- Annotation-Free One-Shot Imitation Learning for Multi-Step Manipulation Tasks
- The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- PRIVMARK: Private Large Language Models Watermarking with MPC
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- EOE: Evolutionary Optimization of Experts for Training Language Models
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- UI-UG: A Unified MLLM for UI Understanding and Generation
- MAS2: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
- SCI-Verifier: Scientific Verifier with Thinking
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
- MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis
- DepthLM: Metric Depth From Vision Language Models
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- BPMN Assistant: An LLM-Based Approach to Business Process Modeling
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Calibrating Verbalized Confidence with Self-Generated Distractors
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- Incentive-Aligned Multi-Source LLM Summaries
- Reinforcement Mid-Training
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
- AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
- HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
- Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
- Learning to Reason in Structured In-context Environments with Reinforcement Learning
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
- The Matthew Effect of AI Programming Assistants: A Hidden Bias in Software Evolution
- C2GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Infusing Theory of Mind into Socially Intelligent LLM Agents
- HEART: Emotionally-driven test-time scaling of Language Models
- Hilbert: Recursively Building Formal Proofs with Informal Reasoning
- WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
- WoW: Towards a World omniscient World model Through Embodied Interaction
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards
- Variational Reasoning for Language Models
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
- A model of errors in transformers
- What Is The Political Content in LLMs' Pre- and Post-Training Data?
- InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning
- Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks
- ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity
- UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
- Polysemous Language Gaussian Splatting via Matching-based Mask Lifting
- Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
- From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
- Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding
- KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
- Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- ReviewScore: Misinformed Peer Review Detection with Large Language Models
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
- GeoEvolve: Automating Geospatial Model Discovery via Multi-Agent Large Language Models
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- Best-of-∞ -- Asymptotic Performance of Test-Time Compute
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- AOT*: Efficient Synthesis Planning via LLM-Empowered AND-OR Tree Search
- StyleBench: Evaluating thinking styles in Large Language Models
- MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
- How Large Language Models Need Symbolism
- ThinkFake: Reasoning in Multimodal Large Language Models for AI-Generated Image Detection
- Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- Logics-Parsing Technical Report
- MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
- The Conductor and the Engine: A Path Towards Co-Designed Reasoning
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- Are Foundation Models Ready for Industrial Defect Recognition? A Reality Check on Real-World Data
- Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
- VIG-RL: Learning to Search and Insert for Verified Image Grounding
- Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
- Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
- Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
- VETO: Towards Protecting Images From Frontier AI Editing
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
- OPENXRD: a comprehensive benchmark framework for LLM/MLLM XRD question answering
- Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
- Benchmark for Assessing Olfactory Perception of Large Language Models
- SleepLM: Natural-Language Intelligence for Human Sleep
- When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
- F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
- NGRPO: Negative-enhanced Group Relative Policy Optimization
- VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
- Spacer: Towards Engineered Scientific Inspiration
- Introducing LongCat-Flash-Thinking: A Technical Report
- StereoFoley: Object-Aware Stereo Audio Generation from Video
- ClassMind: Scaling Classroom Observation and Instructional Feedback with Multimodal AI
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- Qwen3-Omni Technical Report
- Governing Automated Strategic Intelligence
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
- RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- The Universal Personalizer: Few-Shot Dysarthric Speech Recognition via Meta-Learning
- PersonaMatrix: A Recipe for Persona-Aware Evaluation of Legal Summarization
- Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories
- Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- Chain-of-Thought Re-ranking for Image Retrieval Tasks
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
- Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- M-PACE: Mother Child Framework for Multimodal Compliance
- DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
- Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
- FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
- CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusion
- Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
- Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Towards General Agentic Intelligence via Environment Scaling
- The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
- DaSAThco: Data-Aware SAT Heuristics Combinations Optimization via Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- Igniting VLMs toward the Embodied Space
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- AKCIT-FN at CheckThat! 2025: Switching Fine-Tuned SLMs and LLM Prompting for Multilingual Claim Normalization
- VideoAgent: Personalized Synthesis of Scientific Videos
- Towards Better Health Conversations: The Benefits of Context-seeking
- F4-ITS: Fine-grained Feature Fusion for Food Image-Text Search
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- AdsQA: Towards Advertisement Video Understanding
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
- Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference
- Autonomous Code Evolution Meets NP-Completeness
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
- Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
- Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
- Scaling Performance of Large Language Model Pretraining
- ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
- Guideline-Consistent Segmentation via Multi-Agent Refinement
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages
- SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis
- PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
- On Entropy Control in LLM-RL Algorithms
- Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- DCPO: Dynamic Clipping Policy Optimization
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- ShortageSim: Simulating Drug Shortages under Information Asymmetry
- Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- Dream-Coder 7B: An Open Diffusion Language Model for Code
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages
- LongCat-Flash Technical Report
- Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Open Data Synthesis For Deep Research
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
- AWorld: Orchestrating the Training Recipe for Agentic AI
- NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- GDS Agent for Graph Algorithmic Reasoning
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
- Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
- Self-Rewarding Vision-Language Model via Reasoning Decomposition
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- Do MLLMs Really Understand the Charts?
- SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
- Understanding Tool-Integrated Reasoning
- GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- DiscussLLM: Teaching Large Language Models When to Speak
- SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
- Image-Conditioned 3D Gaussian Splat Quantization
- Building and Measuring Trust between Large Language Models
- Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
- RynnEC: Bringing MLLMs into Embodied World
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
- AI Agents for Photonic Integrated Circuit Design Automation
- "DIVE" into Hydrogen Storage Materials Discovery with AI Agents
- Involuntary Jailbreak: On Self-Prompting Attacks
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality
- LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
- Region-Level Context-Aware Multimodal Understanding
- EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
- Benchmarking LLM-based Agents for Single-cell Omics Analysis
- Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- AIM-Bench: Evaluating Decision-making Biases of Agentic LLM as Inventory Manager
- MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- Ovis2.5 Technical Report
- Modeling Human Responses to Multimodal AI Content
- EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems
- Data-Driven Discovery of Interpretable Kalman Filter Variants through Large Language Models and Genetic Programming
- User-centric Subjective Leaderboard by Customizable Reward Modeling
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge
- The Human-AI Hybrid Delphi Model: A Structured Framework for Context-Rich, Expert Consensus in Complex Domains
- Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
- Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
- Aryabhata: An exam-focused language model for JEE Math
- Retrospective Sparse Attention for Efficient Long-Context Generation
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
- Interpreting Fedspeak with Confidence: A LLM-Based Uncertainty-Aware Framework Guided by Monetary Policy Transmission Paths
- Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
- MCPToolBench++: A Large Scale AI Agent Model Context Protocol MCP Tool Use Benchmark
- Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
- EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis
- Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy
- DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- A Neurosymbolic Framework for Interpretable Cognitive Attack Detection in Augmented Reality
- Simulating Human-Like Learning Dynamics with LLM-Empowered Agents
- Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- Operationalizing Serendipity: Multi-Agent AI Workflows for Enhanced Materials Characterization with Theory-in-the-Loop
- StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
- PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- Are Today's LLMs Ready to Explain Well-Being Concepts?
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
- Compressing Chain-of-Thought in LLMs via Step Entropy
- CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
- Principle-Guided Verilog Optimization: IP-Safe Knowledge Transfer via Local-Cloud Collaboration
- CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction
- ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts
- Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
- When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
- Can LLMs Generate High-Quality Task-Specific Conversations?
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- Dynamic Context Adaptation for Consistent Role-Playing Agents with Retrieval-Augmented Generations
- Fine-Tuning Vision-Language Models for Markdown Conversion of Financial Tables in Malaysian Audited Financial Reports
- PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
- TextQuests: How Good are LLMs at Text-Based Video Games?
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
- Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
- CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records
- BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
Discussions
- Google’s Gemini 2.5 paper has 3295 authors arxiv.org/abs/2507.06261 [bsky, 58 points, 7 comments]
- logging off of bluesky to feed my other infinite internet addiction (arxiv) i uh, didn't read the gemini 2.5 report previously so arxiv.org/abs/2507.06261 [bsky, 27 points, 1 comments]
- Gemini 2.5 paper has > 3000 authors arxiv.org/abs/2507.06261 [bsky, 3 points, 0 comments]
- Gemini 2.5 Paper: Advanced Reasoning, Multimodality, Long Context, Agents [hn, 2 points, 1 comments]
- ASK HN: Why Google's Gemini 2.5 paper has 3295 authors? [hn, 2 points, 4 comments]
- arxiv.org/abs/2507.06261 [bsky, 1 points, 0 comments]
- #MLSky - direct link to the paper: arxiv.org/abs/2507.06261 [bsky, 0 points, 0 comments]
- Вот это я понимаю, масштаб! (3195 additional authors not shown) arxiv.org/abs/2507.06261 [bsky, 0 points, 0 comments]
Related