GPT-4o System Card
2024/10/25 by OpenAI, A. M. Hurst, : +503 · 1138 citations
Medicine · #Cardiovascular Function and Risk Factors #Hyperglycemia and glycemic control in critically ill and hospitalized patients
paper · pdf · doi:10.48550/arxiv.2410.21276
Abstract
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning all inputs and outputs are processed by the same neural network. GPT-4o can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time in conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models. In line with our commitment to building AI safely and consistent with our voluntary commitments to the White House, we are sharing the GPT-4o System Card, which includes our Preparedness Framework evaluations. In this System Card, we provide a detailed look at GPT-4o's capabilities, limitations, and safety evaluations across multiple categories, focusing on speech-to-speech while also evaluating text and image capabilities, and measures we've implemented to ensure the model is safe and aligned. We also include third-party assessments on dangerous capabilities, as well as discussion of potential societal impacts of GPT-4o's text and vision capabilities.
Cited by
- ProGuard: Towards Proactive Multimodal Safeguard
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
- Scaling GUI Agents with Visual State Transitions
- Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- PathSelect: Sequential Token Selection for Whole Slide Pathology
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- TopoFE: topology-aware LLM-guided Automated Feature Engineering
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
- VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
- In-Context Learning as Implicit Policy Gradient
- AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
- FARM: Find Anything using Relational Spatial Memory
- CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
- QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction
- Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
- Can an Actor-Critic Optimization Framework Improve Analog Design?
- MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
- Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models
- Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
- Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
- Streaming Video Instruction Tuning
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
- AutoBaxBuilder: Bootstrapping Code Security Benchmarking
- Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies
- Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation
- MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
- Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
- M3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Event Extraction in Large Language Model
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
- TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
- Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
- Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
- X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
- An Agentic AI Framework for Training General Practitioner Student Skills
- VeruSAGE: A Study of Agent-Based Verification for Rust Systems
- Sophia: A Persistent Agent Framework of Artificial Life
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- CheXPO-v2: Preference Optimization for Chest X-ray VLMs with Knowledge Graph Consistency
- LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Generative Refocusing: Flexible Defocus Control from a Single Image
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- Kling-Omni Technical Report
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
- In Pursuit of Pixel Supervision for Visual Pre-training
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- Isolated Sign Language Recognition with Segmentation and Pose Estimation
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
- Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Spherical Leech Quantization for Visual Tokenization and Generation
- ART: Articulated Reconstruction Transformer
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- Scaling Laws for Code: Every Programming Language Matters
- Security and Detectability Analysis of Unicode Text Watermarking Methods Against Large Language Models
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
- TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
- Using GUI Agent for Electronic Design Automation
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Benchmarking the Generality of Vision-Language-Action Models
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- Unveiling User Perceptions in the Generative AI Era: A Sentiment-Driven Evaluation of AI Educational Apps' Role in Digital Transformation of e-Teaching
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
- IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
- Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- LISN: Language-Instructed Social Navigation with VLM-based Controller Modulating
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- Defining Cost Function of Steganography with Large Language Models
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Food Image Generation on Multi-Noun Categories
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- Self-Evolving 3D Scene Generation from a Single Image
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
- Relational Visual Similarity
- Towards Stable Cross-Domain Depression Recognition under Missing Modalities
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
- START: Spatial and Textual Learning for Chart Understanding
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- The Effect of Belief Boxes and Open-mindedness on Persuasion
- LLM Harms: A Taxonomy and Discussion
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
- Training Multi-Image Vision Agents via End2End Reinforcement Learning
- 2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency
- Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- ResearchArcade: Graph Interface for Academic Tasks
- Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- Verbalizing LLMs' assumptions to explain and control sycophancy
- OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
- Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild
- Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
- When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models
- CRAFT-E: A Neuro-Symbolic Framework for Embodied Affordance Grounding
- PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design
- Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance
- GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- Procedural Mistake Detection via Action Effect Modeling
- AsymPuzl: An Asymmetric Puzzle for multi-agent cooperation
- ViDiC: Video Difference Captioning
- SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
- OneThinker: All-in-one Reasoning Model for Image and Video
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- CREST: Universal Safety Guardrails Through Cluster-Guided Cross-Lingual Transfer
- GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Guardian: Detecting Robotic Planning and Execution Errors with Vision-Language Models
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- Evaluating SAM2 for Video Semantic Segmentation
- HiconAgent: History Context-aware Policy Optimization for GUI Agents
- FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing
- Seeing the Wind from a Falling Leaf
- AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
- FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting
- DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
- Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
- SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding
- TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
- CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions
- PhotoFramer: Multi-modal Image Composition Instruction
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- GreenPlanner: Practical Floorplan Layout Generation via an Energy-Aware and Function-Feasible Generative Framework
- When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
- PAT3D: Physics-Augmented Text-to-3D Scene Generation
- Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Geometrically-Constrained Agent for Spatial Reasoning
- ABounD: Adversarial Boundary-Driven Few-Shot Learning for Multi-Class Anomaly Detection
- SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation
- From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
- Vision-Language Memory for Spatial Reasoning
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- Learning Multi-Access Point Coordination in Agentic AI Wi-Fi with Large Language Models
- VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization
- Fara-7B: An Efficient Agentic Model for Computer Use
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- Automating Deception: Scalable Multi-Turn LLM Jailbreaks
- MAGMA-Edu: Multi-Agent Generative Multimodal Framework for Text-Diagram Educational Question Generation
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
- ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
- SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
- TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
- Do Vision-Language Models Understand Visual Persuasiveness?
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Personalized Reward Modeling for Text-to-Image Generation
- TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
- Think Visually, Reason Textually: Vision-Language Synergy in ARC
- First Frame Is the Place to Go for Video Content Customization
- AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
- Computer-Use Agents as Judges for Generative User Interface
- Automatic Pruning Discovery for Large Language Models
- As If We've Met Before: LLMs Exhibit Certainty in Recognizing Seen Files
- EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
- UniSER: A Foundation Model for Unified Soft Effects Removal
- Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
- Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
- DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
- Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- Insight-A: Attribution-aware for Multimodal Misinformation Detection
- Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
- Spatial Blind Spot: Auditory Motion Perception Deficits in Audio LLMs
- Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
- PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
- VEIL: Jailbreaking Text-to-Video Models via Visual Exploitation from Implicit Language
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- Direct Visual Grounding by Directing Attention of Visual Tokens
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- FINRS: A Risk-Sensitive Trading Framework for Real Financial Markets
- Fast Reasoning Segmentation for Images and Videos
- Explainable AI-Generated Image Detection RewardBench
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- AI Agent-Driven Framework for Automated Product Knowledge Graph Construction in E-Commerce
- DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition
- Chain-of-Generation: Progressive Latent Diffusion for Text-Guided Molecular Design
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
- How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
- High-throughput iNaturalist image analysis reveals flower color divergence in <i>Monarda fistulosa</i>
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
- Exploring the Underwater World Segmentation without Extra Training
- LLM-Powered Fully Automated Chaos Engineering: Towards Enabling Anyone to Build Resilient Software Systems at Low Cost
- UniREditBench: A Unified Reasoning-based Image Editing Benchmark
- Quantification and object perception in Multimodal Large Language Models and human linguistic cognition
- DIMO: Diverse 3D Motion Generation for Arbitrary Objects
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
- MuonAll: Muon Variant for Efficient Finetuning of Large Language Models
- Route Experts by Sequence, not by Token
- Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
- A Multi-Agent System for Semantic Mapping of Relational Data to Knowledge Graphs
- When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- AI-assisted workflow enables rapid, high-fidelity breast cancer clinical trial eligibility prescreening
- Culture in Action: Evaluating Text-to-Image Models through Social Activities
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation
- Generating Software Architecture Description from Source Code using Reverse Engineering and Large Language Model
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
- A benchmark multimodal oro-dental dataset for large vision-language models
- Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- DeepEyesV2: Toward Agentic Multimodal Model
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Seeing Straight: Document Orientation Detection for Efficient OCR
- PETRA: Pretrained Evolutionary Transformer for SARS-CoV-2 Mutation Prediction
- GUI-360^∘: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- LLM-as-a-Judge is Bad, Based on AI Attempting the Exam Qualifying for the Member of the Polish National Board of Appeal
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
- Scaling Agent Learning via Experience Synthesis
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- To See or To Read: User Behavior Reasoning in Multimodal LLMs
- Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
- Web-Scale Collection of Video Data for 4D Animal Reconstruction
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- From the Rock Floor to the Cloud: A Systematic Survey of State-of-the-Art NLP in Battery Life Cycle
- A Step Toward World Models: A Survey on Robotic Manipulation
- ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models
- ChartAB: A Benchmark for Chart Grounding & Dense Alignment
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
- Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
- OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
- Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4
- Instance-Level Composed Image Retrieval
- MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- GPTOpt: Towards Efficient LLM-Based Black-Box Optimization
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
- Investigating the Validity Evidence of Automated Scoring Methods for Divergent Thinking Assessments
- Reward Models are Metrics in a Trench Coat
- Do Methods Support the Claims? Intra-Paper Verification for Peer Review
- VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
- CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
- CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
- BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence
- TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
- Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
- EA3D: Online Open-World 3D Object Extraction from Streaming Videos
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- What Limits Agentic Systems Efficiency?
- Quantum Combinatorial Reasoning for Large Language Models
- Uncovering Gaps Between RFC Updates and TCP/IP Implementations: LLM-Facilitated Differential Checks on Intermediate Representations
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
- TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- ChessQA: Evaluating Large Language Models for Chess Understanding
- World Simulation with Video Foundation Models for Physical AI
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- ReCode: Unify Plan and Action for Universal Granularity Control
- ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
- Code Aesthetics with Agentic Reward Feedback
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
- REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
- Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
- Efficient Low Rank Attention for Long-Context Inference in Large Language Models
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- LightAgent: Mobile Agentic Foundation Models
- ArchISMiner: A Framework for Automatic Mining of Architectural Issue-Solution Pairs from Online Developer Communities
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning
- Pctx: Tokenizing Personalized Context for Generative Recommendation
- Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- Social Simulations with Large Language Model Risk Utopian Illusion
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Soft Instruction De-escalation Defense
- SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
- HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
- StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
- Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications
- SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
- Using Large Language Models for Abstraction of Planning Domains - Extended Version
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
- On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- AutoMT: A Multi-Agent LLM Framework for Automated Metamorphic Testing of Autonomous Driving Systems
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- Monitoring LLM-based Multi-Agent Systems Against Corruptions via Node Evaluation
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Prompt fidelity of ChatGPT4o / Dall-E3 text-to-image visualisations
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- Event-Grounding Graph: Unified Spatio-Temporal Scene Graph from Robotic Observations
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- Sherlock Your Queries: Learning to Ask the Right Questions for Dialogue-Based Retrieval
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- Motion2Meaning: A Clinician-Centered Framework for Contestable LLM in Parkinson's Disease Gait Interpretation
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- HouseTour: A Virtual Real Estate A(I)gent
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Disparities in Multilingual LLM-Based Healthcare Q&A
- Token-Level Inference-Time Alignment for Vision-Language Models
- Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Robustness in Text-Attributed Graph Learning: Insights, Trade-offs, and New Defenses
- MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
- TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- Video Reasoning without Training
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- End-to-end Listen, Look, Speak and Act
- SolverLLM: Leveraging Test-Time Scaling for Optimization Problem via LLM-Guided Search
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- Fault Cause Identification across Manufacturing Lines through Ontology-Guided and Process-Aware FMEA Graph Learning with LLMs
- DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
- Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- Benchmarking Multimodal Large Language Models for Face Recognition
- Programmatic Representation Learning with Language Models
- Free-Grained Hierarchical Recognition
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Exploring Cross-Modal Flows for Few-Shot Learning
- Caruca: Effective and Efficient Specification Mining for Opaque Software Components
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Toward Cybersecurity-Expert Small Language Models
- NOSA: Native and Offloadable Sparse Attention
- Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- Generative Universal Verifier as Multimodal Meta-Reasoner
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- Reflection-Based Task Adaptation for Self-Improving VLA
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- Diff-XYZ: A Benchmark for Evaluating Diff Understanding
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
- HoneyBee: Data Recipes for Vision-Language Reasoners
- EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- Inferring Dynamic Physical Properties from Video Foundation Models
- EvoCAD: Evolutionary CAD Code Generation with Vision Language Models
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- Information-Preserving Reformulation of Reasoning Traces for Antidistillation
- mmWalk: Towards Multi-modal Multi-view Walking Assistance
- DocReward: A Document Reward Model for Structuring and Stylizing
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
- Ensembling Large Language Models to Characterize Affective Dynamics in Student-AI Tutor Dialogues
- A Survey on Agentic Multimodal Large Language Models
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
- Embodiment in multimodal large language models
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- Don't Just Fine-tune the Agent, Tune the Environment
- LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
- Synthetic Sandbox for Training Machine Learning Engineering Agents
- The Geometry of Reasoning: Flowing Logics in Representation Space
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- A Systematic Study on Generating Web Vulnerability Proof-of-Concepts Using Large Language Models
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- AutoPR: Let's Automate Your Academic Promotion!
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- LLM Based Long Code Translation using Identifier Replacement
- Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
- ENLighten: Lighten the Transformer, Enable Efficient Optical Acceleration
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
- From Keywords to Clusters: AI-Driven Analysis of YouTube Comments to Reveal Election Issue Salience in 2024
- Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
- Ctrl-VI: Controllable Video Synthesis via Variational Inference
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
- Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents
- Reinforcing Diffusion Models by Direct Group Preference Optimization
- AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents
- Information Seeking for Robust Decision Making under Partial Observability
- DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- OpenJAI-v1.0: An Open Thai Large Language Model
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- SanDRA: Safe Large-Language-Model-Based Decision Making for Automated Vehicles Using Reachability Analysis
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- Towards Better Optimization For Listwise Preference in Diffusion Models
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect
- YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
- Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
- Generative AI-Driven Hierarchical Multi-Agent Framework for Zero-Touch Optical Networks
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- Character Mixing for Video Generation
- Finish First, Perfect Later: Test-Time Token-Level Cross-Validation for Diffusion Large Language Models
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
- Learning to Generate Rigid Body Interactions with Video Diffusion Models
- ContextNav: Towards Agentic Multimodal In-Context Learning
- TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- Where Did It All Go Wrong? A Hierarchical Look into Multi-Agent Error Attribution
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Using predefined vector systems as latent space configuration for neural network supervised training on data with arbitrarily large number of classes
- Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs
- H-DDx: A Hierarchical Evaluation Framework for Differential Diagnosis
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- Cross-Modal Content Optimization for Steering Web Agent Preferences
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
- AudioToolAgent: An Agentic Framework for Audio-Language Models
- Reward Model Routing in Alignment
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
- Learning Efficient Guardrails for Compliance
- EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- Syntax-Guided Diffusion Language Models with User-Integrated Personalization
- QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL
- Graph-S3: Enhancing Agentic textual Graph Retrieval with Synthetic Stepwise Supervision
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- Hearing the Order: Investigating Selection Bias in Large Audio-Language Models
- When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- Exploring System 1 and 2 communication for latent reasoning in LLMs
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
- Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
- MuSLR: Multimodal Symbolic Logical Reasoning
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- DescribeEarth: Describe Anything for Remote Sensing Images
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- IRIS: Intrinsic Reward Image Synthesis
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
- Scaling Synthetic Task Generation for Agents via Exploration
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Diamonds in the rough: Transforming SPARCs of imagination into a game concept by leveraging medium sized LLMs
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- Understanding the Dilemma of Unlearning for Large Language Models
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
- When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
- Prompt and Parameter Co-Optimization for Large Language Models
- MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis
- Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- BPMN Assistant: An LLM-Based Approach to Business Process Modeling
- An AI-guided framework for automated map point symbol generation through template rendering
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- PhysiAgent: An Embodied Agent Framework in Physical World
- Expanding Computation Spaces of LLMs at Inference Time
- Query Circuits: Explaining How Language Models Answer User Prompts
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
- HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based Fuzzing
- Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Diff-3DCap: Shape Captioning with Diffusion Models
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- Toward a Holistic Approach to Continual Model Merging
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow
- Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
- Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
- The Matthew Effect of AI Programming Assistants: A Hidden Bias in Software Evolution
- Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
- Robot Learning from Any Images
- WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
- The price of automated case law annotation: comparing the cost and performance of GPT-4o and student annotators
- Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
- Text Adversarial Attacks with Dynamic Outputs
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
- InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
- The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
- Developing Vision-Language-Action Model from Egocentric Videos
- RISK: A Framework for GUI Agents in E-commerce Risk Management
- Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
- Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- RobustFlow: Towards Robust Agentic Workflow Generation
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Plan2Evolve: LLM Self-Evolution for Improved Planning Capability via Automated Domain Generation
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- AOT*: Efficient Synthesis Planning via LLM-Empowered AND-OR Tree Search
- Look Before you Leap: Estimating LLM Benchmark Scores from Descriptions
- Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems
- Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning
- Thinking Augmented Pre-training
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- Responsible AI Technical Report
- Embodied AI: From LLMs to World Models
- DiffNator: Generating Structured Explanations of Time-Series Differences
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
- Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
- DP-LENS: A Density-Aware Polyfocal Lens with Topology-Driven Auto-Routing for Occlusion Management in Immersive 3D Analytics
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
- Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
- Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
- OPENXRD: a comprehensive benchmark framework for LLM/MLLM XRD question answering
- References Improve LLM Alignment in Non-Verifiable Domains
- RepoTransAgent: Multi-Agent LLM Framework for Repository-Aware Code Translation
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
- OSDA: A Framework for Open-Set Discovery and Automatic Interpretation of Land-cover in Remote Sensing Imagery
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Benchmarking PDF Accessibility Evaluation A Dataset and Framework for Assessing Automated and LLM-Based Approaches for Accessibility Testing
- Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- A closed-loop AI framework for hypothesis-driven and interpretable materials design
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
- Enhancing the NAO: Extending Capabilities of Legacy Robots for Long-Term Research
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
- AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
- Mano Technical Report
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- Evaluating Generative AI as an Educational Tool for Radiology Resident Report Drafting
- ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification
- Variation in Verification: Understanding Verification Dynamics in Large Language Models
- Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
- Similarity Field Theory: A Mathematical Framework for Intelligence
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- Stencil: Subject-Driven Generation with Context Guidance
- AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
- AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
- LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation
- A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- AISTAT lab system for DCASE2025 Task6: Language-based audio retrieval
- Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
- LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- DISCO: Disentangled Communication Steering for Large Language Models
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
- On Optimal Steering to Achieve Exact Fairness
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
- Chain-of-Thought Re-ranking for Image Retrieval Tasks
- From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
- SimCoachCorpus: A naturalistic dataset with language and trajectories for embodied teaching
- Assessing Historical Structural Oppression Worldwide via Rule-Guided Prompting of Large Language Models
- CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
- Exploring the Capabilities of LLM Encoders for Image-Text Retrieval in Chest X-rays
- MICA: Multi-Agent Industrial Coordination Assistant
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
- Synthetic Data Generation for Screen Time and App Usage
- Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- The Few-shot Dilemma: Over-prompting Large Language Models
- OVITA: Open-Vocabulary Interpretable Trajectory Adaptations
- Multimodal Hate Detection Using Dual-Stream Graph Neural Networks
- Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
- LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- Pluralistic Off-policy Evaluation and Alignment
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
- VideoAgent: Personalized Synthesis of Scientific Videos
- DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
- ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
- Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results
- EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
- BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Multimodal LLMs See Sentiment
- AdsQA: Towards Advertisement Video Understanding
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- The Carbon Footprint Wizard: A Knowledge-Augmented AI Interface for Streamlining Food Carbon Footprint Analysis
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions
- Unleashing the True Potential of LLMs: A Feedback-Triggered Self-Correction with Long-Term Multipath Decoding
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
- Automated Hierarchical Graph Construction for Multi-source Electronic Health Records
- ALPHA: LLM-Enabled Active Learning for Human-Free Network Anomaly Detection
- Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
- UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
- TAGAL: Tabular Data Generation using Agentic LLM Methods
- Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
- PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- The Impact of Critique on LLM-Based Model Generation from Natural Language: The Case of Activity Diagrams
- ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Plan More, Debug Less: Applying Metacognitive Theory to AI-Assisted Programming Education
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
- Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- Fidelity-preserving enhancement of ptychography with foundational text-to-image models
- Kwai Keye-VL 1.5 Technical Report
- Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions
- TopoNav: Topological Graphs as a Key Enabler for Advanced Object Navigation
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics
- A Survey on Current Trends and Recent Advances in Text Anonymization
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
- Efficient Code Embeddings from Code Generation Models
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
- CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models
- AvatarBack: Back-Head Generation for Complete 3D Avatars from Front-View Images
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- AWorld: Orchestrating the Training Recipe for Agentic AI
- End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Learning to Generate Unit Test via Adversarial Reinforcement Learning
- GRAFT: GRaPH and Table Reasoning for Textual Alignment -- A Benchmark for Structured Instruction Following and Visual Reasoning
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- Evaluating Language Model Reasoning about Confidential Information
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- The Art of Hide and Seek: Making Pickle-Based Model Supply Chain Poisoning Stealthy Again
- Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- LongReasonArena: A Long Reasoning Benchmark for Large Language Models
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Reasoning LLMs in the Medical Domain: A Literature Survey
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph Learning
- Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
- Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Building and Measuring Trust between Large Language Models
- Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis
- Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
- Generics and Default Reasoning in Large Language Models
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- CRED-SQL: Enhancing Real-world Large Scale Database Text-to-SQL Parsing through Cluster Retrieval and Execution Description
- Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery
- Leveraging Large Language Models for Predictive Analysis of Human Misery
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality
- Region-Level Context-Aware Multimodal Understanding
- Meet Your New Client: Writing Reports for AI -- Benchmarking Information Loss in Market Research Deliverables
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- MAD: A Benchmark for Multi-Turn Audio Dialogue Fact-Checking
- Scalable RF Simulation in Generative 4D Worlds
- Talk Less, Fly Lighter: Autonomous Semantic Compression for UAV Swarm Communication via LLMs
- Benchmarking LLM-based Agents for Single-cell Omics Analysis
- LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
- ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
- FairTabGen: Unifying Counterfactual and Causal Fairness in Synthetic Tabular Data Generation
- Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- Agentic Design Review System
- ORBIT: An Object Property Reasoning Benchmark for Visual Inference Tasks
- Bridging Solidity Evolution Gaps: An LLM-Enhanced Approach for Smart Contract Compilation Error Resolution
- Evaluating LLMs on Chinese Idiom Translation
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
- The Roots of International Perceptions: Simulating US Attitude Changes Towards China with LLM Agents
- DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
- From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
- Leveraging Large Language Models for Rare Disease Named Entity Recognition
- Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
- Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
- SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling
- LL3M: Large Language 3D Modelers
- Large Language Models for Subjective Language Understanding: A Survey
- Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
- What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction
- DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
- An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance
- Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text Guidance
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- "Draw me a curator" Examining the visual stereotyping of a cultural services profession by generative AI
- EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis
- Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs
- VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- gpt-oss-120b & gpt-oss-20b Model Card
- LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
- PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- LinguaFluid: Language Guided Fluid Control via Semantic Rewards in Reinforcement Learning
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Valid Inference with Imperfect Synthetic Data
- In-Training Defenses against Emergent Misalignment in Language Models
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- Do Machines Think Emotionally? Cognitive Appraisal Analysis of Large Language Models
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- Single-Step Reconstruction-Free Anomaly Detection and Segmentation via Diffusion Models
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- From MAS to MARS: Coordination Failures and Reasoning Trade-offs in Hierarchical Multi-Agent Robotic Systems within a Healthcare Scenario
- Query Attribute Modeling: Improving search relevance with Semantic Search and Meta Data Filtering
- Training-Free Multimodal Large Language Model Orchestration
- SID: Benchmarking Guided Instruction Capabilities in STEM Education with a Socratic Interdisciplinary Dialogues Dataset
- ESDD 2026: Environmental Sound Deepfake Detection Challenge Evaluation Plan
- Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
- Automated Generation of Curriculum-Aligned Multiple-Choice Questions for Malaysian Secondary Mathematics Using Generative AI
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- HarmonyGuard: Toward Safety and Utility in Web Agents via Adaptive Policy Enhancement and Dual-Objective Optimization
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Characterizing Deep Research: A Benchmark and Formal Definition
- DTPA: Dynamic Token-level Prefix Augmentation for Controllable Text Generation
- VeriGUI: Verifiable Long-Chain GUI Dataset
- Sotopia-RL: Reward Design for Social Intelligence
- OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
- R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
- Estimating Worst-Case Frontier Risks of Open-Weight LLMs
- FFHQ-Makeup: Paired Synthetic Makeup Dataset with Facial Consistency Across Multiple Styles
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
- Data Dependency-Aware Code Generation from Enhanced UML Sequence Diagrams
- MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
- ReflecSched: Solving Dynamic Flexible Job-Shop Scheduling via LLM-Powered Hierarchical Reflection
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy
- Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
- Evading Data Provenance in Deep Neural Networks
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple Judges
- PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
- Accurate and Consistent Graph Model Generation from Text with Large Language Models
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
Related