Gemma 3 Technical Report
2025/03/25 by Gemma Team, Aishwarya Kamath, Kamath, Aishwarya +418 · 3 voices · 462 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2503.19786
Abstract
We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision understanding abilities, a wider coverage of languages and longer context - at least 128K tokens. We also change the architecture of the model to reduce the KV-cache memory that tends to explode with long context. This is achieved by increasing the ratio of local to global attention layers, and keeping the span on local attention short. The Gemma 3 models are trained with distillation and achieve superior performance to Gemma 2 for both pre-trained and instruction finetuned versions. In particular, our novel post-training recipe significantly improves the math, chat, instruction-following and multilingual abilities, making Gemma3-4B-IT competitive with Gemma2-27B-IT and Gemma3-27B-IT comparable to Gemini-1.5-Pro across benchmarks. We release all our models to the community.
Cited by
- A Structured Clustering Approach for Inducing Media Narratives
- The Appeal and Reality of Recycling LoRAs with Adaptive Merging
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
- The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
- With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- Towards Understanding Steering Strength
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- Leveraging Semantic Maps for City-Scale Cross-View Localization
- FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
- Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
- Detect Before You Leap: Mirage Detection in Vision-Language Models
- EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
- An Information Theoretic Perspective on Agentic System Design
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- Can abstract concepts from LLM improve SLM performance?
- Shuttling Compiler for Trapped-Ion Quantum Computers Based on Large Language Models
- Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
- Xiaomi MiMo-VL-Miloco Technical Report
- EasyV2V: A High-quality Instruction-based Video Editing Framework
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- Sigma-MoE-Tiny Technical Report
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- DSO: Direct Steering Optimization for Bias Mitigation
- The Deleuzian Representation Hypothesis
- Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess Evaluation
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- MedCEG: Reinforcing Verifiable Medical Reasoning with Critical Evidence Graph
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org
- Leveraging LLMs for Title and Abstract Screening for Systematic Review: A Cost-Effective Dynamic Few-Shot Learning Approach
- ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
- CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- LASER: Language Model Regression for Semi-Structured Workflow Resource and Runtime Estimation
- Generalized Referring Expression Segmentation on Aerial Photos
- FOAM: Blocked State Folding for Memory-Efficient LLM Training
- RunawayEvil: Jailbreaking the Image-to-Video Generative Models
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
- Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models
- Personalizing Agent Privacy Decisions via Logical Entailment
- Generative AI Practices, Literacy, and Divides: An Empirical Analysis in the Italian Context
- M3DR: Towards Universal Multilingual Multimodal Document Retrieval
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- Menta: A Small Language Model for On-Device Mental Health Prediction
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- Multilingual Training-Free Remote Sensing Image Captioning
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
- Aligning Probabilistic Beliefs under Informative Missingness: LLM Steerability in Clinical Reasoning
- When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- Optimizing Multimodal Language Models through Attention-based Interpretability
- Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
- Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
- Explainable Visual Anomaly Detection via Concept Bottleneck Models
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- CREward: A Type-Specific Creativity Reward Model
- AppSelectBench: Application-Level Tool Selection Benchmark
- Wrist Photoplethysmography Predicts Dietary Information
- Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- Vidi2: Large Multimodal Models for Video Understanding and Creation
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Do Vision-Language Models Understand Visual Persuasiveness?
- Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
- NLP Datasets for Idiom and Figurative Language Tasks
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- NeuroPath: Neurobiology-Inspired Path Tracking and Reflection for Semantically Coherent Retrieval
- Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
- Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
- SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
- Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
- How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
- STORM: Segment, Track, and Object Re-Localization from a Single Image
- State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Voice-Interactive Surgical Agent for Multimodal Patient Data Control
- More Agents Helps but Adversarial Robustness Gap Persists
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
- Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
- S2LM: Towards Semantic Steganography via Large Language Models
- Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
- Motif 2 12.7B technical report
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- Context informs pragmatic interpretation in vision-language models
- TabGemma: Text-Based Tabular ICL via LLM using Continued Pretraining and Retrieval
- AutoHood3D: A Multi-Modal Benchmark for Automotive Hood Design and Fluid-Structure Interaction
- Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- Advancing Subsurface Discovery and Geothermal Monitoring with an Agentic Artificial Intelligence Framework
- How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
- MiRAGE: Misconception Detection with Retrieval-Guided Multi-Stage Reasoning and Ensemble Fusion
- Learning When to Quit in Sales Conversations
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- Dynamic Reflections: Probing Video Representations with Text Alignment
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
- Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
- Toward Sustainability-Aware LLM Inference on Edge Clusters
- Proactive DDoS Detection and Mitigation in Decentralized Software-Defined Networking via Port-Level Monitoring and Zero-Training Large Language Models
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?
- ParaScopes: What do Language Models Activations Encode About Future Text?
- Diffusion LLMs are Natural Adversaries for any LLM
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaboration
- MisSynth: Improving MISSCI Logical Fallacies Classification with Synthetic Data
- RCScore: Quantifying Response Consistency in Large Language Models
- MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
- Alibaba International E-commerce Product Search Competition DcuRAGONs Team Technical Report
- Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- Multimodal Language Models Cannot Spot Spatial Inconsistencies
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
- SCOUT: A Lightweight Framework for Scenario Coverage Assessment in Autonomous Driving
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- zFLoRA: Zero-Latency Fused Low-Rank Adapters
- Conflict Adaptation in Vision-Language Models
- ChessQA: Evaluating Large Language Models for Chess Understanding
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering
- Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
- Model-Aware Tokenizer Transfer
- S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- Head Pursuit: Probing Attention Specialization in Multimodal Transformers
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Epipolar Geometry Improves Video Generation Models
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- ComProScanner: A multi-agent based framework for composition-property structured data extraction from scientific literature
- What Does It Take to Build a Performant Selective Classifier?
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- Data-Centric Lessons To Improve Speech-Language Pretraining
- TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- Vision-language models learn the geometry of human perceptual space
- Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
- The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- Automated HIV Screening on Dutch Electronic Health Records with Large Language Models
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning
- Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
- KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
- Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
- BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
- BeLLMan: Controlling LLM Congestion
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
- Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- FraQAT: Quantization Aware Training with Fractional bits
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Toward Cybersecurity-Expert Small Language Models
- Training LLM Agents to Empower Humans
- LLM one-shot style transfer for Authorship Attribution and Verification
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
- Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
- Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
- ACADATA: Parallel Dataset of Academic Data for Machine Translation
- Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
- A Survey on Agentic Multimodal Large Language Models
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- LightSAE: Parameter-Efficient and Heterogeneity-Aware Embedding for IoT Multivariate Time Series Forecasting
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
- bio.tools: an expanded web service for research software in the life sciences
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- How do LLMs Compute Verbal Confidence
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- On the Entity-Level Alignment in Crosslingual Consistency
- Stability of Transformers under Layer Normalization
- Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
- Gold Panning: Turning Positional Bias into Signal for Multi-Document LLM Reasoning
- PatentVision: A multimodal method for drafting patent applications
- Cross-Platform Narrative Prediction: Leveraging Platform-Invariant Discourse Networks
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- ICL-Router: In-Context Learned Model Representations for LLM Routing
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- A Large-scale Dataset for Robust Complex Anime Scene Text Detection
- MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
- Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations
- Mid-Training of Large Language Models: A Survey
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- CodeVaani: A Multilingual, Voice-Based Code Learning Assistant
- Primal-Dual Direct Preference Optimization for Constrained LLM Alignment
- Adversarial Reinforcement Learning for Large Language Model Agent Safety
- Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
- Robustness assessment of large audio language models in multiple-choice evaluation
- Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding
- Evaluating Self-Supervised Speech Models via Text-Based LLMS
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Distilling Reasoning into Student LLMs: Local Naturalness for Selecting Teacher Data
- Quantitative Certification of Agentic Tool Selection
- What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models
- PsycholexTherapy: Simulating Reasoning in Psychotherapy with Small Language Models in Persian
- H-DDx: A Hierarchical Evaluation Framework for Differential Diagnosis
- MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Learning Efficient Guardrails for Compliance
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- Feature Identification via the Empirical NTK
- Making, not Taking, the Best of N
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- ACT: Agentic Classification Tree
- TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- Training Agents Inside of Scalable World Models
- Expanding Computation Spaces of LLMs at Inference Time
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- Sequential Diffusion Language Models
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- NanoFlux: Adversarial Dual-LLM Evaluation and Distillation For Multi-Domain Reasoning
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
- JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
- HEART: Emotionally-driven test-time scaling of Language Models
- Partial Parameter Updates for Efficient Distributed Training
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
- Multilingual Vision-Language Models, A Survey
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- Code once, Run Green: Automated Green Code Translation in Serverless Computing
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- Training-Free Multimodal Deepfake Detection via Graph Reasoning
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Plan2Evolve: LLM Self-Evolution for Improved Planning Capability via Automated Domain Generation
- On Code-Induced Reasoning in LLMs
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- StyleBench: Evaluating thinking styles in Large Language Models
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Mamba Modulation: On the Length Generalization of Mamba
- Semantic-Aware Fuzzing: An Empirical Framework for LLM-Guided, Reasoning-Driven Input Mutation
- SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
- Retrieval Feedback Memory Enhancement Large Model Retrieval Generation Method
- MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
- Investigating Traffic Accident Detection Using Multimodal Large Language Models
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
- Agentic Reasoning for Robust Vision Systems via Increased Test-Time Compute
- Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
- Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
- Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
- Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- The Few-shot Dilemma: Over-prompting Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- Structural Damage Detection Using AI Super Resolution and Visual Language Model
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Biomedical Hypothesis Explainability with Graph-Based Context Retrieval
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- F4-ITS: Fine-grained Feature Fusion for Food Image-Text Search
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Measuring Epistemic Humility in Multimodal Large Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla
- Quality Assessment of Tabular Data using Large Language Models and Code Generation
- PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Text-Trained LLMs Can Zero-Shot Extrapolate PDE Dynamics, Revealing a Three-Stage In-Context Learning Mechanism
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- TemporalFlowViz: Parameter-Aware Visual Analytics for Interpreting Scramjet Combustion Evolution
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Guideline-Consistent Segmentation via Multi-Agent Refinement
- Transition Models: Rethinking the Generative Learning Objective
- IPA: An Information-Reconstructive Input Projection Framework for Efficient Foundation Model Adaptation
- Binary Quantization For LLMs Through Dynamic Grouping
- CEQuest: Benchmarking Large Language Models for Construction Estimation
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
- Just-in-time and distributed task representations in language models
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- Logical Reasoning with Outcome Reward Models for Test-Time Scaling
- SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- ReflectivePrompt: Reflective evolution in autoprompting algorithms
- The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation
- Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
- Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation
- Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
- LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
- Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
- SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
- From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
- Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Perturbed Public Voices (P2V): A Dataset for Robust Audio Deepfake Detection
- Scaling Up Active Testing to Large Language Models
- Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
- RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
- Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
- Effortless Vision-Language Model Specialization in Histopathology without Annotation
- CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
- Grounding Multilingual Multimodal LLMs With Cultural Knowledge
- HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
- Deep Language Geometry: Constructing a Metric Space from LLM Weights
- MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
- Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
- MV-Debate: Multi-view Agent Debate with Dynamic Reflection Gating for Multimodal Harmful Content Detection in Social Media
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
- Operationalizing Serendipity: Multi-Agent AI Workflows for Enhanced Materials Characterization with Theory-in-the-Loop
- Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- VQA support to Arabic Language Learning Educational Tool
- ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts
- SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
- Semantic Structure in Large Language Model Embeddings
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- Isolating Culture Neurons in Multilingual Large Language Models
- Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
- CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
- Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
- TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection
- Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
- Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework
- NeedleChain: Measuring Intact Long-Context Reasoning Capability of Large Language Models
Discussions
Related