LLaMA: Open and Efficient Foundation Language Models
2023/02/27 by Hugo Touvron, Touvron, Hugo, Thibaut Lavril +25 · 1660 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2302.13971
Abstract
We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B. We release all our models to the research community.
Cited by
- CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- Understanding the Mechanisms of Fast Hyperparameter Transfer
- Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
- Reservoir Computing inspired Matrix Multiplication-free Language Model
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Argus: Token Aware Distributed LLM Inference Optimization
- CNSight: Evaluation of Clinical Note Segmentation Tools
- Visual Autoregressive Modelling for Monocular Depth Estimation
- M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation
- Learning When Not to Attend Globally
- EmoCtrl: Controllable Emotional Image Content Generation
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- LLMBoost: Make Large Language Models Stronger with Boosting
- Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
- MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion
- Attention Residuals
- MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
- Training Language Models to Cooperate with Inference-Time Controllers
- Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation
- SymStep: Symbolic Step Verification for Logical Reasoning
- Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming
- Grevo: A Unified Generative Recommendation Framework with Evolutionary Item Indexing
- SeedFold: Scaling Biomolecular Structure Prediction
- LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests
- Real-time Reconstruction of Human Visual Perception from fMRI
- PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- Context as a Tool: Context Management for Long-Horizon SWE-Agents
- CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
- ARdena: Scenario-driven control of real-time LLM agents
- TokenMem: Faithful Knowledge Injection for Frozen LLMs
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
- StegaFFD: Privacy-preserving Face Forgery Detection via Fine-grained Steganographic Domain Lifting
- Compressing LLMs with MoP: Mixture of Pruners
- PixelGen: Improving Pixel Diffusion with Perceptual Supervision
- GQ-VAE: A gated quantized VAE for learning variable length tokens
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
- A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
- Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
- RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- Input-Adaptive Visual Preprocessing for Efficient Fast Vision-Language Model Inference
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- Benchmarking LLMs for Predictive Applications in the Intensive Care Units
- Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
- Adaptive Financial Sentiment Analysis for NIFTY 50 via Instruction-Tuned LLMs , RAG and Reinforcement Learning Approaches
- QuarkAudio Technical Report
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
- Vehicle-centric Perception via Multimodal Structured Pre-training
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
- HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection
- SAP: Syntactic Attention Pruning for Transformer-based Language Models
- HyperLoad: A Cross-Modality Enhanced Large Language Model-Based Framework for Green Data Center Cooling Load Prediction
- R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
- Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
- Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- A Theoretical Lens for RL-Tuned Language Models via Energy-Based Models
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Phoneme-based speech recognition driven by large language models and sampling marginalization
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
- Holistic Evaluation of State-of-the-Art LLMs for Code Generation
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Adversarial Robustness of Vision in Open Foundation Models
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- Easy Adaptation: An Efficient Task-Specific Knowledge Injection Method for Large Models in Resource-Constrained Environments
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Vision-Language Model Guided Image Restoration
- Subjective Question Generation and Answer Evaluation using NLP
- LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- A systematic assessment of Large Language Models for constructing two-level fractional factorial designs
- LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Radiology Report Generation with Layer-Wise Anatomical Attention
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- CKA-Guided Modular Quantization: Beyond Bit-Width to Algorithmic Diversity
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Sigma-MoE-Tiny Technical Report
- Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
- The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
- An Information-Theoretic Framework for Robust Large Language Model Editing
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
- Sceniris: A Fast Procedural Scene Generation Framework
- Cross-Language Bias Examination in Large Language Models
- Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- In-Context Semi-Supervised Learning
- Large Video Planner Enables Generalizable Robot Control
- FAME: Fictional Actors for Multilingual Erasure
- Null-LoRA: Low-Rank Adaptation on Null Space
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- An Exploratory Study of Bayesian Prompt Optimization for Test-Driven Code Generation with Large Language Models
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- Georeferencing complex relative locality descriptions with large language models
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
- CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- Toward Agentic Environments: GenAI and the Convergence of AI, Sustainability, and Human-Centric Spaces
- Do Reviews Matter for Recommendations in the Era of Large Language Models?
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- FROC: A Unified Framework with Risk-Optimized Control for Machine Unlearning in LLMs
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
- Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
- Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping
- SPAR: Session-based Pipeline for Adaptive Retrieval on Legacy File Systems
- PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- Low-Rank Compression of Language Models via Differentiable Rank Selection
- Supervised Contrastive Frame Aggregation for Video Representation Learning
- WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
- Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
- Fully Inductive Node Representation Learning via Graph View Transformation
- CADMorph: Geometry-Driven Parametric CAD Editing via a Plan-Generate-Verify Loop
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis
- Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
- BAID: A Benchmark for Bias Assessment of AI Detectors
- CLINIC: Evaluating Multilingual Trustworthiness in Language Models for Healthcare
- Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
- CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
- Stronger Normalization-Free Transformers
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
- ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
- The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- Scaling Behavior of Discrete Diffusion Language Models
- DynaMate: An Autonomous Agent for Protein-Ligand Molecular Dynamics Simulations
- DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- Chasing Shadows: Pitfalls in LLM Security Research
- View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
- Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- WhatsCode: Large-Scale GenAI Deployment for Developer Efficiency at WhatsApp
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- Llama-based source code vulnerability detection: Prompt engineering vs Fine tuning
- NeurIDA: Dynamic Modeling for Effective In-Database Analytics
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
- Chat with UAV -- Human-UAV Interaction Based on Large Language Models
- PolyLingua: Margin-based Inter-class Transformer for Robust Cross-domain Language Detection
- CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- Flash Multi-Head Feed-Forward Network
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- M-STAR: Multi-Scale Spatiotemporal Autoregression for Human Mobility Modeling
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing
- VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
- Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- Large Language Model-Based Generation of Discharge Summaries
- Statistic-Augmented, Decoupled MoE Routing and Aggregating in Autonomous Driving
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Small Language Models Can Use Nuanced Reasoning For Health Science Research Classification: A Microbial-Oncogenesis Case Study
- Teaching Language Models Mechanistic Explainability Through Arrow-Pushing
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Convergence of Outputs When Two Large Language Models Interact in a Multi-Agentic Setup
- Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
- Empathy by Design: Aligning Large Language Models for Healthcare Dialogue
- BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
- Compass: Co-Exploration of Mapping and Hardware for Heterogeneous Multi-Chiplet Accelerators Targeting LLM Inference Service Workloads
- LLM Harms: A Taxonomy and Discussion
- HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
- Big Tech-Funded AI Papers Have Higher Citation Impact, Greater Insularity, and Larger Recency Bias
- Fast SceneScript: Accurate and Efficient Structured Language Model via Multi-Token Prediction
- Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems
- Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Text Rationalization for Robust Causal Effect Estimation
- David vs. Goliath: Can Small Models Win Big with Agentic AI in Hardware Design?
- LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
- Bridging Traditional Machine Learning and Large Language Models: A Two-Part Course Design for Modern AI Education
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- Language Models as Semantic Teachers: Post-Training Alignment for Medical Audio Understanding
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- DeRA: Decoupled Representation Alignment for Video Tokenization
- Solving LLM Repetition Problem in Production: A Comprehensive Study of Multiple Solutions
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- NAS-LoRA: Empowering Parameter-Efficient Fine-Tuning for Visual Foundation Models with Searchable Adaptation
- YOLOA: Real-Time Affordance Detection via LLM Adapter
- Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis
- CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models
- Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
- TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- PopSim: Social Network Simulation for Social Media Popularity Prediction
- Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- VACoT: Rethinking Visual Data Augmentation with VLMs
- promptolution: A Unified, Modular Framework for Prompt Optimization
- Decentralized Multi-Agent System with Trust-Aware Communication
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
- Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- Low-Rank Prehab: Preparing Neural Networks for SVD Compression
- Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation
- FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting
- DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
- Financial Instruction Following Evaluation (FIFE)
- M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- FOM-Nav: Frontier-Object Maps for Object Goal Navigation
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
- IRPO: Boosting Image Restoration via Post-training GRPO
- Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
- Preventing Model Collapse via Contraction-Conditioned Neural Filters
- ProEx: A Unified Framework Leveraging Large Language Model with Profile Extrapolation for Recommendation
- Financial Text Classification Based On rLoRA Finetuning On Qwen3-8B model
- SelfAI: Building a Self-Training AI System with LLM Agents
- Breaking It Down: Domain-Aware Semantic Segmentation for Retrieval Augmented Generation
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Teleportation-Based Defenses for Privacy in Approximate Machine Unlearning
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- LUMOS: Large User MOdels for User Behavior Prediction
- Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
- A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
- Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- AI/ML Model Cards in Edge AI Cyberinfrastructure: towards Agentic AI
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- How to Correctly Report LLM-as-a-Judge Evaluations
- GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
- CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
- Graph-O1 : Monte Carlo Tree Search with Reinforcement Learning for Text-Attributed Graph Reasoning
- Towards a Foundation Model for Partial Differential Equations Across Physics Domains
- EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics
- AudioScene: Integrating Object-Event Audio into 3D Scenes
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
- Operator Learning at Machine Precision
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
- A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- In-Context Compositional Learning via Sparse Coding Transformer
- Terminal Velocity Matching
- Gender Bias in Emotion Recognition by Large Language Models
- LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
- Robot-Powered Data Flywheels: Deploying Robots in the Wild for Continual Data Collection and Foundation Model Adaptation
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- LLMAID: Identifying AI Capabilities in Android Apps with LLMs
- IRSDA: An Agent-Orchestrated Framework for Enterprise Intrusion Response
- Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- A Systematic Study of Compression Ordering for Large Language Models
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- UFO: Unfair-to-Fair Evolving Mitigates Unfairness in LLM-based Recommender Systems via Self-Play Fine-tuning
- Enhancing Large Language Models for Automated Homework Assessment in Undergraduate Circuit Analysis
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- ArtiWorld: LLM-Driven Articulation of 3D Objects in Scenes
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- Plan-X: Instruct Video Generation via Semantic Planning
- Towards Efficient LLM-aware Heterogeneous Graph Learning
- EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants
- PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- The Rapid Growth of AI Foundation Model Usage in Science
- Selective Rotary Position Embedding
- Large Language Models for Sentiment Analysis to Detect Social Challenges: A Use Case with South African Languages
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- Geometric-disentangelment Unlearning
- Spanning Tree Autoregressive Visual Generation
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
- Predicting one-year clinical instability and mortality in heart failure patients using sequence modeling
- AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
- ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
- SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
- Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
- On 10x Better Scalability: KV Stores Scale Up KV Cache
- ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
- Boosting Medical Visual Understanding From Multi-Granular Language Learning
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- LLaVA3: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
- Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
- PeerCoPilot: A Language Model-Powered Assistant for Behavioral Health Organizations
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
- Standardising the NLP Workflow: A Framework for Reproducible Linguistic Analysis
- Automatic Pruning Discovery for Large Language Models
- Parameter Importance-Driven Continual Learning for Foundation Models
- PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- UniSER: A Foundation Model for Unified Soft Effects Removal
- What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning
- Strategic Innovation Management in the Age of Large Language Models Market Intelligence, Adaptive R&D, and Ethical Governance
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- Gradient-descent methods for quantum detector tomography
- Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning
- MiAD: Mirage Atom Diffusion for De Novo Crystal Generation
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- 10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
- Knowledge-Grounded Agentic Large Language Models for Multi-Hazard Understanding from Reconnaissance Reports
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright Protection
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Performance and interpretability analysis of code generation large language models
- Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- Medical Knowledge Intervention Prompt Tuning for Medical Image Classification
- Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior?
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Adaptive Focus Memory for Language Models
- MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless Sensing
- Suppressing VLM Hallucinations with Spectral Representation Filtering
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
- GraphPilot: Grounded Scene Graph Conditioning for Language-Based Autonomous Driving
- Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots
- POTSA: A Cross-Lingual Speech Alignment Framework for Low Resource Speech-to-Text Translation
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- Mastering Olympiad-Level Physics with Artificial Intelligence
- From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
- iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation
- The Path Not Taken: RLVR Provably Learns Off the Principals
- The Open Syndrome Definition as a Machine-Readable Standard for Public Health: Design and Implementation Study
- A mathematical theory of balancing relational generalization and memorization
- General Intelligence-based Fragmentation (GIF): A framework for peak-labeled spectra simulation
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- Mitigating Negative Flips via Margin Preserving Training
- Adaptive Multi-Agent Response Refinement in Conversational Systems
- Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
- Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection
- Testing Question Answering Software with Context-Driven Question Generation
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
- Endpoint Security Agent: A Comprehensive Approach to Real-time System Monitoring and Threat Detection
- Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
- Cortex AISQL: A Production SQL Engine for Unstructured Data
- Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
- Importance-Aware Data Selection for Efficient LLM Instruction Tuning
- DIMO: Diverse 3D Motion Generation for Arbitrary Objects
- Voice-Interactive Surgical Agent for Multimodal Patient Data Control
- Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis
- TuckA: Hierarchical Compact Tensor Experts for Efficient Fine-Tuning
- P3-LLM: An Integrated NPU-PIM Accelerator for LLM Inference Using Hybrid Numerical Formats
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention
- AUTO-Explorer: Automated Data Collection for GUI Agent
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- BuildingWorld: A Structured 3D Building Dataset for Urban Foundation Models
- ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning
- LLM3-DTI: A Large Language Model and Multi-modal data co-powered framework for Drug-Target Interaction prediction
- 10 Open Challenges Steering the Future of Vision-Language-Action Models
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- A Remarkably Efficient Paradigm to Multimodal Large Language Models for Sequential Recommendation
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- Adaptation and Fine-tuning with TabPFN for Travelling Salesman Problem
- MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
- S2LM: Towards Semantic Steganography via Large Language Models
- Order-Level Attention Similarity Across Language Models: A Latent Commonality
- Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image Representation
- Scientific judgment drifts over time in AI ideation
- A benchmark multimodal oro-dental dataset for large vision-language models
- The Future of Fully Homomorphic Encryption System: from a Storage I/O Perspective
- BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
- Quantifying the Climate Risk of Generative AI: Region-Aware Carbon Accounting with G-TRACE and the AI Sustainability Pyramid
- Cambrian-S: Towards Spatial Supersensing in Video
- ScaleDL: Towards Scalable and Efficient Runtime Prediction for Distributed Deep Learning Workloads
- Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
- GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
- IndicSuperTokenizer: An Optimized Tokenizer for Indic Multilingual LLMs
- Learning When to Quit in Sales Conversations
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- Credit Network Modeling and Analysis via Large Language Models
- Dynamic Reflections: Probing Video Representations with Text Alignment
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
- Next Token Knowledge Tracing: Exploiting Pretrained LLM Representations to Decode Student Behaviour
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
- An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
- Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
- BRAINS: A Retrieval-Augmented System for Alzheimer's Detection and Monitoring
- Random Initialization of Gated Sparse Adapters
- LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
- GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
- HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
- Assessing Factual Music Comprehension in Large Audio Language Models
- Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
- AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
- Automated Invoice Data Extraction: Using LLM and OCR
- DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
- Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus
- Proactive DDoS Detection and Mitigation in Decentralized Software-Defined Networking via Port-Level Monitoring and Zero-Training Large Language Models
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection
- Data-Efficient Domain Adaptation for LLM-based MT using Contrastive Preference Optimization
- Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- Synergistic Tensor and Pipeline Parallelism
- Identifying the Periodicity of Information in Natural Language
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Elastic Architecture Search for Efficient Language Models
- The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
- LLMs Process Lists With General Filter Heads
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
- Evontree: Ontology Rule-Guided Self-Evolution of Large Language Models
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- Chain-of-Thought Hijacking
- UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
- MisSynth: Improving MISSCI Logical Fallacies Classification with Synthetic Data
- RCScore: Quantifying Response Consistency in Large Language Models
- Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
- QuantumBench: A Benchmark for Quantum Problem Solving
- Predicate Renaming via Large Language Models
- Revisiting Multilingual Data Mixtures in Language Model Pretraining
- SoK: Honeypots & LLMs, More Than the Sum of Their Parts?
- π_
RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models - Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
- CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions
- SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
- ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
- Factuality challenges in the era of large language models and opportunities for fact-checking
- Beyond the “Wow” Factor: Using Generative AI for Increasing Generative Sense-Making
- AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
- Emission-Forecasting-Based Spatial-Temporal Carbon Response: A Multi-Agent Attention-Enhanced Deep Learning Framework
- LLM-FP4: 4-Bit Floating-Point Quantized Transformers
- Mixture-of-Depths Attention
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
- ExplainRec: Towards Explainable Multi-Modal Zero-Shot Recommendation with Preference Attribution and Large Language Models
- Self-driving laboratories in Japan
- Exploring large language models for the generation of synthetic training samples for aspect-based sentiment analysis in low resource settings
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- RWGBench: Evaluating Scholarly Positioning in Related Work Generation
- Flows: Building Blocks of Reasoning and Collaborating AI
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- Why Open Source? A Game-Theoretic Analysis of the AI Race
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- SENSE: Efficient EEG-to-Text via Privacy-Preserving Semantic Retrieval
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
- RL makes MLLMs see better than SFT
- MIN-Merging: Merge the Important Neurons for Model Merging
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- Revisiting scalable sequential recommendation with Multi-Embedding Approach and Mixture-of-Experts
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
- Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
- Suzume-chan: Your Personal Navigator as an Embodied Information Hub
- Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- What Limits Agentic Systems Efficiency?
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
- zFLoRA: Zero-Latency Fused Low-Rank Adapters
- Synergizing chemical and AI communities for advancing laboratories of the future
- SALS: Sparse Attention in Latent Space for KV cache Compression
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Leveraging LLMs for Early Alzheimer's Prediction
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
- ChessQA: Evaluating Large Language Models for Chess Understanding
- Hybrid Modeling, Sim-to-Real Reinforcement Learning, and Large Language Model Driven Control for Digital Twins
- Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Larger and more instructable language models become less reliable
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- AutoStreamPipe: LLM Assisted Automatic Generation of Data Stream Processing Pipelines
- EMTSF:Extraordinary Mixture of SOTA Models for Time Series Forecasting
- MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models
- Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs
- Fast-MIA: Efficient and Scalable Membership Inference for LLMs
- Can Language Models Compose Skills In-Context?
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Normalization in Attention Dynamics
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- A Unified Model for Multi-Task Drone Routing in Post-Disaster Road Assessment
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- Efficient semantic uncertainty quantification in language models via diversity-steered sampling
- MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection
- Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Understanding Network Behaviors through Natural Language Question-Answering
- LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
- Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs
- UniSE: A Unified Framework for Decoder-only Autoregressive LM-based Speech Enhancement
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- On Interaction Effects in Greybox Fuzzing
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Review of Tools for Zero-Code LLM Based Application Development
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
- CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
- KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- Restoring Pruned Large Language Models via Lost Component Compensation
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Semantic World Models
- SecureInfer: Heterogeneous TEE-GPU Architecture for Privacy-Critical Tensors for Large Language Model Deployment
- RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
- C2T-ID: Converting Semantic Codebooks to Textual Document Identifiers for Generative Search
- XGen-Q: An Explainable Domain-Adaptive LLM Framework with Retrieval-Augmented Generation for Software Security
- QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models
- ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
- Large language models in medicine
- Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- SegTune: Structured and Fine-Grained Control for Song Generation
- BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
- Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
- NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
- Identity-Aware Large Language Models require Cultural Reasoning
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Planned Diffusion
- MEG-GPT: A transformer-based foundation model for magnetoencephalography data
- Language Models as Semantic Augmenters for Sequential Recommenders
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Using natural language processing to analyse text data in behavioural science
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Strengthening LLMs for Tabular Prediction with Structural Priors
- TaxoAlign: Scholarly Taxonomy Generation Using Language Models
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
- Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
- Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Graph4MM: Weaving Multimodal Learning with Structural Information
- Back to Bytes: Revisiting Tokenization Through UTF-8
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
- Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
- Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
- QuanBench: Benchmarking Quantum Code Generation with Large Language Models
- MuonBP: Faster Muon via Block-Periodic Orthogonalization
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Prompt Optimization via Retrieved Reasoning Assets and Multi-Agent Analysis
- Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
- Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
- Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
- Learning to Answer from Correct Demonstrations
- The 3rd Place Solution of CCIR CUP 2025: A Framework for Retrieval-Augmented Generation in Multi-Turn Legal Conversation
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
- Attention Is All You Need for KV Cache in Diffusion LLMs
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- Scaling Artificial Intelligence for Multi-Tumor Early Detection with More Reports, Fewer Masks
- Cross-Scenario Unified Modeling of User Interests at Billion Scale
- Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
- Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
- Cognitive-Aligned Spatio-Temporal Large Language Models For Next Point-of-Interest Prediction
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
- Generalist vs Specialist Time Series Foundation Models: Investigating Potential Emergent Behaviors in Assessing Human Health Using PPG Signals
- IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
- Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training
- David vs. Goliath: A comparative study of different-sized LLMs for code generation in the domain of automotive scenario generation
- Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
- MADREC: A Multi-Aspect Driven LLM Agent for Explainable and Adaptive Recommendation
- BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- NeuroRVQ: Multi-Scale EEG Tokenization for Generative Large Brainwave Models
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- DSCD: Large Language Model Detoxification with Self-Constrained Decoding
- RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- A Survey on Evaluation of Large Language Models
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- Traveling Salesman-Based Token Ordering Improves Stability in Homomorphically Encrypted Language Models
- A Survey on Parallel Reasoning
- Single chip 1 Tb/s optical transmitter with inverse designed input and output couplers
- Evaluating the Quality of Randomness and Entropy in Tasks Supported by Large Language Models
- Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
- Indoor Localization using Compact, Telemetry-Agnostic, Transfer-Learning Enabled Decoder-Only Transformer
- ADARL: Adaptive Low-Rank Structures for Robust Policy Learning under Uncertainty
- Scaling Language-Centric Omnimodal Representation Learning
- BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
- MeTA-LoRA: Data-Efficient Multi-Task Fine-Tuning for Large Language Models
- Efficient LLM Inference over Heterogeneous Edge Networks with Speculative Decoding
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- FedLoRA-Optimizer: Federated LoRA Fine-Tuning with Global and Local Optimization in Heterogeneous Data Scenarios
- XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
- CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis
- FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
- Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
- Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- Scalable and Explainable Enterprise Knowledge Discovery Using Graph-Centric Hybrid Retrieval
- Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
- PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
- PAGE: Prompt Augmentation for text Generation Enhancement
- Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition
- Direct Multi-Token Decoding
- Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- High-Fidelity Speech Enhancement via Discrete Audio Tokens
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication
- Large Language Model-Empowered Channel Prediction and Predictive Beamforming for LEO Satellite Communications
- Enhancing Long-Term Welfare in Recommender Systems: An Information Revelation Approach
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- SoundReactor: Frame-level Online Video-to-Audio Generation
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
- Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- Embodiment in multimodal large language models
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- Unpacking Hateful Memes: Presupposed Context and False Claims
- Shaping History, Responsibly: Seven Principles to Guide the Design of Archival AI Assistants for Cultural Heritage Collections
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
- Exploring Large Language Models for Financial Applications: Techniques, Performance, and Challenges with FinMA
- Learning a Dense Reasoning Reward Model from Expert Demonstration via Inverse Reinforcement Learning
- Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
- Detecting LLM-Generated Spam Reviews by Integrating Language Model Embeddings and Graph Neural Network
- OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
- Accelerating Attention with Basis Decomposition
- Small edits, large models: How Wikipedia advocacy shapes LLM values
- Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
- Holistic Order Prediction in Natural Scenes
- Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
- Can We Reliably Rank Model Performance across Domains without Labeled Data?
- StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
- Provable Watermarking for Data Poisoning Attacks
- SynthID-Image: Image watermarking at internet scale
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- AdaPM: a Partial Momentum Algorithm for LLM Training
- Efficient Resource-Constrained Training of Vision Transformers via Subspace Optimization
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Preference-Aware Memory Update for Long-Term LLM Agents
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Web Crawler Restrictions, AI Training Datasets & Political Biases
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
- Constraints-of-Thought: A Framework for Constrained Reasoning in Language-Model-Guided Search
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Provable Training Data Identification for Large Language Models
- When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach
- GRETEL: A Goal-driven Retrieval and Execution-based Trial Framework for LLM Tool Selection Enhancing
- PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- Agent Learning via Early Experience
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
- Instance Relation Learning Network with Label Knowledge Propagation for Few-shot Multi-label Intent Detection
- Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention
- OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Detecting Post-generation Edits to Watermarked LLM Outputs via Combinatorial Watermarking
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- Learning to Decide with Just Enough: Information-Theoretic Context Summarization for CMDPs
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Mid-Training of Large Language Models: A Survey
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
- Reasoning by Exploration: A Unified Approach to Retrieval and Generation over Graphs
- The Framework That Survives Bad Models: Human-AI Collaboration For Clinical Trials
- Training Dynamics Impact Post-Training Quantization Robustness
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
- Mixture of Neuron Experts
- Primal-Dual Direct Preference Optimization for Constrained LLM Alignment
- Membership Inference Attacks on Tokenizers of Large Language Models
- YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
- vAttention: Verified Sparse Attention
- AutoPentester: An LLM Agent-based Framework for Automated Pentesting
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- Prototype-Based Dynamic Steering for Large Language Models
- Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
- Operational Risks in Grid Integration of Large Data Center Loads: Characteristics, Stability Assessments, and Sensitivity Studies
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
- Self Speculative Decoding for Diffusion Large Language Models
- Evaluation of Clinical Trials Reporting Quality using Large Language Models
- The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
- Small Language Models for Emergency Departments Decision Support: A Benchmark Study
- Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
- Efficient Test-Time Scaling for Small Vision-Language Models
- Know Thyself? On the Incapability and Implications of AI Self-Recognition
- IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
- TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
- Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Geolog-IA: Conversational System for Academic Theses
- Deep Generative Continual Learning using Functional LoRA: FunLoRA
- Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
- MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
- Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
- Backdoor Attacks Against Speech Language Models
- Integrating AI and Ensemble Forecasting: Explainable Materials Planning with Scorecards and Trend Insights for a Large-Scale Manufacturer
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Visual Self-Refinement for Autoregressive Models
- The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- Automated Structured Radiology Report Generation with Rich Clinical Context
- SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
- Generative AI for subgrid turbulence in large-eddy simulations
- TsLLM: Augmenting LLMs for General Time Series Understanding and Prediction
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
- Retrieval-Augmented Framework for LLM-Based Clinical Decision Support
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- Instruction‐Tuned Large‐Language Models for Quality Control in Automatic Item Generation: A Feasibility Study
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
- SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- The silence of the weights: an investigation of structural pruning strategies for attention-based audio signal architectures
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- MUVLA: Learning to Explore Object Navigation via Map Understanding
- ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
- HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
- RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- ScheduleMe: Multi-Agent Calendar Assistant
- Layer-wise dynamic rank for compressing large language models
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- Effective Model Pruning
- OceanGym: A Benchmark Environment for Underwater Embodied Agents
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- EEsizer: LLM-Based AI Agent for Sizing of Analog and Mixed Signal Circuit
- Towards Structured Knowledge: Advancing Triple Extraction from Regional Trade Agreements using Large Language Models
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- Vision Function Layer in Multimodal LLMs
- SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment
- Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
- When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
- VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- NeMo: Needle in a Montage for Video-Language Understanding
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
- Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
- Reinforcement Mid-Training
- Autoregressive Video Generation beyond Next Frames Prediction
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
- Sequential Diffusion Language Models
- The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact
- Detecting and Rectifying Noisy Labels: A Similarity-based Approach
- ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
- Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- Spiral of Silence in Large Language Model Agents
- Aligning LLMs for Multilingual Consistency in Enterprise Applications
- Sparse-Up: Learnable Sparse Upsampling for 3D Generation with High-Fidelity Textures
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- The Impact of Role Design in In-Context Learning for Large Language Models
- GeoBS: Information-Theoretic Quantification of Geographic Bias in AI Models
- NeuroBridge: Using Generative AI to Bridge Cross-neurotype Communication Differences through Neurotypical Perspective-taking
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- Fin-ExBERT: User Intent based Text Extraction in Financial Context using Graph-Augmented BERT and trainable Plugin
- Adaptive Token-Weighted Differential Privacy for LLMs: Not All Tokens Require Equal Protection
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
- Limit Analysis for Symbolic Multi-step Reasoning Tasks with Information Propagation Rules Based on Transformers
- PT2-LLM: Post-Training Ternarization for Large Language Models
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Data-Efficient Training by Evolved Sampling
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
- XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
- d2Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching
- MindCraft: How Concept Trees Take Shape In Deep Models
- MMPB: It's Time for Multi-Modal Personalization
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- Voice user interfaces for effortless navigation in medical virtual reality environments
- Ethische Betrachtungen der automatisierten Textanalyse <b>: Fortschritte, Risiken und die Notwendigkeit eines Gleichgewichts</b>
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Activation Function Design Sustains Plasticity in Continual Learning
- Meta-Awareness Enhances Reasoning Models: Self-Alignment Reinforcement Learning
- Investigating Faithfulness in Large Audio Language Models
- Stochastic activations
- SoDaDE: Solvent Data-Driven Embeddings with Small Transformer Models
- Unlocking the Power of Mixture-of-Experts for Task-Aware Time Series Analytics
- Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
- Multi-Agent Path Finding via Offline RL and LLM Collaboration
- Multilingual Vision-Language Models, A Survey
- Mind the Missing: Variable-Aware Representation Learning for Irregular EHR Time Series using Large Language Models
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- Debiasing Large Language Models in Thai Political Stance Detection via Counterfactual Calibration
- A High-Capacity and Secure Disambiguation Algorithm for Neural Linguistic Steganography
- AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
- KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
- On the Complexity Theory of Masked Discrete Diffusion: From poly(1/ε) to Nearly ε-Free
- Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- Tiny-QMoE
- InvBench: Can LLMs Accelerate Program Verification with Invariant Synthesis?
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- Go With The Flow: Churn-Tolerant Decentralized Training of Large Language Models
- PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
- A short survey on almost orthogonal vectors in a few specific large dimensions
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
- Generative AI for FFRDCs
- ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
- Instruction Boundary: Quantifying Biases in LLM Reasoning under Various Coverage
- One Filters All: A Generalist Filter for State Estimation
- Embodied AI: From LLMs to World Models
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
- Eliminating stability hallucinations in llm-based tts models via attention guidance
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- Learning Contextual Retrieval for Robust Conversational Search
- Boosting Zero-Shot VLN via Abstract Obstacle Map-Based Waypoint Prediction with TopoGraph-and-VisitInfo-Aware Prompting
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
- Uncertainty in Semantic Language Modeling with PIXELS
- epiGPTope: A Machine Learning-Based Epitope Generator and Classifier
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
- VIG-RL: Learning to Search and Insert for Verified Image Grounding
- The potential and limitations of large language models for automatic classification of teachers' motivational messages in educational research
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Divergence Decoding: Training-Free Capability Fusion
- Reading Between the Signs: Predicting Future Suicidal Ideation from Adolescent Social Media Texts
- SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
- Training Compute-Optimal Protein Language Models
- From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
- Chatbot Epistemology
- AI for social science and social science of AI: A survey
- Navigating the AI era: university communication strategies and perspectives on generative AI tools
- Multi-Robot system environmental constraint analysis by petri nets
- Retrieval Feedback Memory Enhancement Large Model Retrieval Generation Method
- VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
- ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
- Enhancing Speech Large Language Models through Reinforced Behavior Alignment
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- MAPO: Mixed Advantage Policy Optimization
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Code Driven Planning with Domain-Adaptive Critic
- Advances in Large Language Models for Medicine
- Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
- ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
- CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
- Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
- Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs
- Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
- AccessEval: Benchmarking Disability Bias in Large Language Models
- How Persuasive is Your Context?
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- LIMI: Less is More for Agency
- nDNA -- the Semantic Helix of Artificial Cognition
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- ACCeLLiuM: Supervised Fine-Tuning for Automated OpenACC Pragma Generation
- Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
- Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- \boldsymbolλ-Orthogonality Regularization for Compatible Representation Learning
- Federated Learning with Ad-hoc Adapter Insertions: The Case of Soft-Embeddings for Training Classifier-as-Retriever
- DoubleGen: Debiased Generative Modeling of Counterfactuals
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- REFER: Mitigating Bias in Opinion Summarisation via Frequency Framed Prompting
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
- Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction
- Defining and Monitoring Complex Robot Activities via LLMs and Symbolic Reasoning
- Efficient Multimodal Dataset Distillation via Generative Models
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
- Benchmarking and Improving LLM Robustness for Personalized Generation
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
- Rationality Check! Benchmarking the Rationality of Large Language Models
- Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
- Adaptive LoRA Experts Allocation and Selection for Federated Fine-Tuning
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- AToken: A Unified Tokenizer for Vision
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- Enhancing Time Awareness in Generative Recommendation
- TFMAdapter: Lightweight Instance-Level Adaptation of Foundation Models for Forecasting with Covariates
- Neural Speech Separation with Parallel Amplitude and Phase Spectrum Estimation
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
- Learning the natural history of human disease with generative transformers
- GestOS: Advanced Hand Gesture Interpretation via Large Language Models to control Any Type of Robot
- Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
- Large Language Model-Empowered Decision Transformer for UAV-Enabled Data Collection
- Benchmarking ChatGPT and DeepSeek in April 2025: A Novel Dual Perspective Sentiment Analysis Using Lexicon-Based and Deep Learning Approaches
- Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
- Validating Solidity Code Defects using Symbolic and Concrete Execution powered by Large Language Models
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- Toward PDDL Planning Copilot
- Sparse Training Scheme for Multimodal LLM
- Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
- Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- Bi-level Personalization for Federated Foundation Models: A Task-vector Aggregation Approach
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- KoSEL: Knowledge subgraph enhanced large language model for medical question answering
- A comparison of pipelines for the translation of a low resource language based on transformers
- RAGs to Riches: RAG-like Few-shot Learning for Large Language Model Role-playing
- Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
- Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training
- Token Homogenization under Positional Bias
- LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
- Uncertainty in Authorship: Why Perfect AI Detection Is Mathematically Impossible
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptation
- A Dynamic Knowledge Update-Driven Model with Large Language Models for Fake News Detection
- AssemMate: Graph-Based LLM for Robotic Assembly Assistance
- ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- Does Language Model Understand Language?
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- BERT4beam: Large AI Model Enabled Generalized Beamforming Optimization
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- Adapting Public Personas: A Multimodal Study of U.S. Legislators' Cross-Platform Social Media Strategies
- CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- Compartmentalised Agentic Reasoning for Clinical NLI
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Large Language Models Meet Legal Artificial Intelligence: A Survey
- Beyond Token Limits: Assessing Language Model Performance on Long Text Classification
- Securing LLM-Generated Embedded Firmware through AI Agent-Driven Validation and Patching
- WALL: A Web Application for Automated Quality Assurance using Large Language Models
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis
- Prompting the Market? A Large-Scale Meta-Analysis of GenAI in Finance NLP (2022-2025)
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- ENSI: Efficient Non-Interactive Secure Inference for Large Language Models
- Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
- LightAgent: Production-level Open-source Agentic AI Framework
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
- GmSLM : Generative Marmoset Spoken Language Modeling
- HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
- Character-Level Perturbations Disrupt LLM Watermarks
- TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla
- Meta-Learning Reinforcement Learning for Crypto-Return Prediction
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
- BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
- Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution
- Multimodal LLMs See Sentiment
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- PICO: Performance Insights for Collective Operations
- TraceRAG: A LLM-Based Framework for Explainable Android Malware Detection and Behavior Analysis
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Bias after Prompting: Persistent Discrimination in Large Language Models
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
- Language Self-Play For Data-Free Training
- Towards Post-mortem Data Management Principles for Generative AI
- AdaMixT: Adaptive Weighted Mixture of Multi-Scale Expert Transformers for Time Series Forecasting
- AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
- Causal Attention with Lookahead Keys
- Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- Large language models surpass domain-specific architectures for antepartum electronic fetal monitoring analysis
- Reconstruction Alignment Improves Unified Multimodal Models
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects
- The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- SoK: Security and Privacy of AI Agents for Blockchain
- UniSearch: Rethinking Search System with a Unified Generative Architecture
- Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments
- Contrastive Self-Supervised Network Intrusion Detection using Augmented Negative Pairs
- FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
- Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
- Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
- Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
- LESER: Learning to Expand via Search Engine-feedback Reinforcement in e-Commerce
- Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- Semantic-guided LoRA Parameters Generation
- Masked Diffusion Language Models with Frequency-Informed Training
- OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Reverse Browser: Vector-Image-to-Code Generator
- A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning
- Towards Open World Detection: A Survey
- Shared Autonomy through LLMs and Reinforcement Learning for Applications to Ship Hull Inspections
- Delta Activations: A Representation for Finetuned Large Language Models
- Towards a Unified View of Large Language Model Post-Training
- Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
- PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- SMooGPT: Stylized Motion Generation using Large Language Models
- SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning
- Weakly-Supervised Learning of Dense Functional Correspondences
- VulRTex: A Reasoning-Guided Approach to Identify Vulnerabilities from Rich-Text Issue Report
- Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
- PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
- LLM-GUARD: Large Language Model-Based Detection and Repair of Bugs and Security Vulnerabilities in C++ and Python
- SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
- PromptCOS: Towards Content-only System Prompt Copyright Auditing for LLMs
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- AIVA: An AI-based Virtual Companion for Emotion-aware Interaction
- TraceLLM: Security Diagnosis Through Traces and Smart Contracts in Ethereum
- Advancing SLM Tool-Use Capability using Reinforcement Learning
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- Generative AI for Crystal Structures: A Review
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- CMRAG: Co-modality-based visual document retrieval and question answering
- From Confidence to Collapse in LLM Factual Robustness
- Better by Comparison: Retrieval-Augmented Contrastive Reasoning for Automatic Prompt Optimization
- HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers
- FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- Deep Reinforcement Learning for Drone Route Optimization in Post-Disaster Road Assessment
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Upcycling Candidate Tokens of Large Language Models for Query Expansion
- 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- Benchmarking and Studying the LLM-based Code Review
- CSRM-LLM: Embracing Multilingual LLMs for Cold-Start Relevance Matching in Emerging E-commerce Markets
- Cloud-Device Collaborative Agents for Sequential Recommendation
- Insight-LLM: LLM-enhanced Multi-view Fusion in Insider Threat Detection
- The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion
- LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
- MARS: Modality-Aligned Retrieval for Sequence Augmented CTR Prediction
- Question-to-Knowledge (Q2K): Multi-Agent Generation of Inspectable Facts for Product Mapping
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- X-Troll: eXplainable Detection of State-Sponsored Information Operations Agents
- TableZoomer: A Collaborative Agent Framework for Large-scale Table Question Answering
- A User-centric Kubernetes-based Architecture for Green Cloud Computing
- Efficient Graph Understanding with LLMs via Structured Context Injection
- OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
- MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- Neuro-Symbolic Predictive Process Monitoring
- Spotlighter: Revisiting Prompt Tuning from a Representative Mining View
- MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Framework
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
- Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
- BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting
- COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
- Entropy-based Coarse and Compressed Semantic Speech Representation Learning
- Universal Properties of Activation Sparsity in Modern Large Language Models
- The Resurgence of GCG Adversarial Attacks on Large Language Models
- Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
- Intelligent Spectrum Management in Satellite Communications
- Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
- Data Auctions for Retrieval Augmented Generation
- Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
- EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
- VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
- Rethinking Layer-wise Model Merging through Chain of Merges
- Generalizable Object Re-Identification via Visual In-Context Prompting
- HyperFlexis: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
- WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
- PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
- Signs of Struggle: Spotting Cognitive Distortions across Language and Register
- Can News Predict the Direction of Oil Price Volatility? A Language Model Approach with SHAP Explanations
- SemSR: Semantics aware robust Session-based Recommendations
- Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
- AWorld: Orchestrating the Training Recipe for Agentic AI
- MPFormer: Adaptive Framework for Industrial Multi-Task Personalized Sequential Retriever
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Foundation Models for Cross-Domain EEG Analysis Application: A Survey
- Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
- Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
- OneRec-V2 Technical Report
- Automated Bug Triaging using Instruction-Tuned Large Language Models
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- Beyond Transcription: Mechanistic Interpretability in ASR
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- Refining Text Generation for Realistic Conversational Recommendation via Direct Preference Optimization
- Self-supervised structured object representation learning
- Ego-centric Predictive Model Conditioned on Hand Trajectories
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- FinCast: A Foundation Model for Financial Time-Series Forecasting
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning
- SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
- Ensemble Debates with Local Large Language Models for AI Alignment
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC
- GENIE-ASI: Generative Instruction and Executable Code for Analog Subcircuit Identification
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees
- An Investigation on Group Query Hallucination Attacks
- MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
- Reflection-Enhanced Meta-Optimization Integrating TextGrad-style Prompt Optimization with Memory-Driven Self-Evolution
- Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
- Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
- VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- RETAIL: Towards Real-world Travel Planning for Large Language Models
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Influence-driven Curriculum Learning for Pre-training on Limited Data
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- Identifying and Answering Questions with False Assumptions: An Interpretable Approach
- Fine-tuning and prompt engineering for large language models-based code review automation
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
- Long-Context Speech Synthesis with Context-Aware Memory
- SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
- Comparing energy consumption and accuracy in text classification inference
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- Online Conformal Selection with Accept-to-Reject Changes
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- MGT-Prism: Enhancing Domain Generalization for Machine-Generated Text Detection via Spectral Alignment
- From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models
- AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- FLAIR: Feedback Learning for Adaptive Information Retrieval
- Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
- Maximum Score Routing For Mixture-of-Experts
- CRED-SQL: Enhancing Real-world Large Scale Database Text-to-SQL Parsing through Cluster Retrieval and Execution Description
- Unlearning Comparator: A Visual Analytics System for Comparative Evaluation of Machine Unlearning Methods
- MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation
- CC-Time: Cross-Model and Cross-Modality Time Series Forecasting
- GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- TBGRecall: A Generative Retrieval Model for E-commerce Recommendation Scenarios
- UniCast: A Unified Framework for Instance-Conditioned Multimodal Time-Series Forecasting
- Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
- AgentMental: An Interactive Multi-Agent Framework for Explainable and Adaptive Mental Health Assessment
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
- Retrieval-augmented reasoning with lean language models
- Feedback Indicators: The Alignment between Llama and a Teacher in Language Learning
- Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
- Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Modeling Human Responses to Multimodal AI Content
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- REFN: A Reinforcement-Learning-From-Network Framework against 1-day/n-day Exploitations
- AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
- SemPT: Semantic Prompt Tuning for Vision-Language Models
- Increasing the Utility of Synthetic Images through Chamfer Guidance
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- Flexible Personalized Split Federated Learning for On-Device Fine-Tuning of Foundation Models
- Beyond Semantic Understanding: Preserving Collaborative Frequency Components in LLM-based Recommendation
- Inductive Bias Extraction and Matching for LLM Prompts
- Failures to Surface Harmful Contents in Video Large Language Models
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- Teaching LLMs to Speak Spectroscopy
- Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- Preacher: Paper-to-Video Agentic System
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Enhancing Memory Recall in LLMs with Gauss-Tin: A Hybrid Instructional and Gaussian Replay Approach
- NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
- DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
- Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring
- DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
- InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
- A Survey on Training-free Alignment of Large Language Models
- LLaMA-Based Models for Aspect-Based Sentiment Analysis
- Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment
- Agentic Graph Neural Networks for Wireless Communications and Networking Towards Edge General Intelligence: A Survey
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- KFFocus: Highlighting Keyframes for Enhanced Video Understanding
- Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- DiffractGPT: Atomic Structure Determination from X-ray Diffraction Patterns using Generative Pre-trained Transformer
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Grouped Speculative Decoding for Autoregressive Image Generation
- LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
- Large Language Models for Subjective Language Understanding: A Survey
- Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
- MolmoAct: Action Reasoning Models that can Reason in Space
- Tailored Emotional LLM-Supporter: Enhancing Cultural Sensitivity
- UrzaGPT: LoRA-Tuned Large Language Models for Card Selection in Collectible Card Games
- Semantic-Enhanced Time-Series Forecasting via Large Language Models
- ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
- AutoAssert 1: A LoRA Fine-Tuned LLM Model for Efficient Automated Assertion Generation
- Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Semantic Alignment
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- QuiZSF: An efficient data-model interaction framework for zero-shot time-series forecasting
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning
- CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
- Bridging Classical and Quantum Computing for Next-Generation Language Models
- Whisfusion: Parallel ASR Decoding with Masked Diffusion
- Confidence Estimation for Text-to-SQL in Large Language Models
- HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning
- Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
- Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
- Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- When a Paper Has 1000 Authors: Rethinking Citation Metrics in the Era of LLMs
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- Learning by Teaching: Engaging Students as Instructors of Large Language Models in Computer Science Education
- Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
- LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- Streamlining Admission with LOR Insights: AI-Based Leadership Assessment in Online Master's Program
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
- AI-assisted JSON Schema Creation and Mapping
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
- ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos
- An Effective Approach for Node Classification in Textual Graphs
- VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
- Multimodal RAG Enhanced Visual Description
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
- X-SAM: From Segment Anything to Any Segmentation
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
- RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake Detection
- SVGen: Interpretable Vector Graphics Generation with Large Language Models
- TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- DRAMA: A Dynamic and Robust Allocation-based Multi-Agent System for Changing Environments
- StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution
- Leveraging large language models for SQL behavior-based database intrusion detection
- From eye to AI: studying rodent social behavior in the era of machine Learning
- DP-GPT4MTS: Dual-Prompt Large Language Model for Textual-Numerical Time Series Forecasting
- Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
- COPO: Consistency-Aware Policy Optimization
- AquaChat++: LLM-Assisted Multi-ROV Inspection for Aquaculture Net Pens with Integrated Battery Management and Thruster Fault Tolerance
- Efficient Scaling for LLM-based ASR
- PAIRS: Parametric-Verified Adaptive Information Retrieval and Selection for Efficient RAG
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency
- RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory
- GP and LLMs for Program Synthesis: No Clear Winners
- Sotopia-RL: Reward Design for Social Intelligence
- An Entity Linking Agent for Question Answering
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
- Can Large Vision-Language Models Understand Multimodal Sarcasm?
- Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
- VQA support to Arabic Language Learning Educational Tool
- Semantic-aware Graph-guided Behavior Sequences Generation with Large Language Models for Smart Homes
- IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
- CTTS: Collective Test-Time Scaling
- LOST: Low-rank and Sparse Pre-training for Large Language Models
- TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
- Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- Evaluating Position Bias in Large Language Model Recommendations
- Toward Efficient Spiking Transformers: Synapse Pruning Meets Synergistic Learning-Based Compensation
- Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving
- SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
- Large Language Model Guided Decoding for Self-Supervised Speech Recognition
- Zero-shot Compositional Action Recognition with Neural Logic Constraints
- GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- Empowering Tabular Data Preparation with Language Models: Why and How?
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- T-GRAG: A Dynamic GraphRAG Framework for Resolving Temporal Conflicts and Redundancy in Knowledge Retrieval
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
- KCR: Resolving Long-Context Knowledge Conflicts via Reasoning in LLMs
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
- CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- SaviorRec: Semantic-Behavior Alignment for Cold-Start Recommendation
- T2S: Tokenized Skill Scaling for Lifelong Imitation Learning
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- A Note on Code Quality Score: LLMs for Maintainable Large Codebases
- How LLMs are Shaping the Future of Virtual Reality
- PaPaformer: Language Model from Pre-trained Parallel Paths
- AutoDebias: Automated Framework for Debiasing Text-to-Image Models
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- ITUNLP at SemEval-2025 Task 8: Question-Answering over Tabular Data: A Zero-Shot Approach using LLM-Driven Code Generation
- Uncovering Latent Connections in Indigenous Heritage: Semantic Pipelines for Cultural Preservation in Brazil
- Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- ART: Adaptive Relation Tuning for Generalized Relation Prediction
- A Unified Perception-Language-Action Framework for Adaptive Autonomous Driving
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- PixNerd: Pixel Neural Field Diffusion
- OKG-LLM: Aligning Ocean Knowledge Graph with Observation Data via LLMs for Global Sea Surface Temperature Prediction
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- Failures Are the Stepping Stones to Success: Enhancing Few-Shot In-Context Learning by Leveraging Negative Samples
- ChatVis: Large Language Model Agent for Generating Scientific Visualizations
- SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
- KLLM: Fast LLM Inference with K-Means Quantization
- Automatically discovering heuristics in a complex SAT solver with large language models
- TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
- Real-time News Story Identification
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- What is Beneath Misogyny: Misogynous Memes Classification and Explanation
- DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
- A Foundation Model for Material Fracture Prediction
- Context-aware Rotary Position Embedding
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- Improving Generative Ad Text on Facebook using Reinforcement Learning
Related