QLoRA: Efficient Finetuning of Quantized LLMs
2023/05/23 by Tim Dettmers, Dettmers, Tim, Artidoro Pagnoni +5 · 3 voices · 508 citations
#cs.LG
paper · pdf · doi:10.48550/arxiv.2305.14314
Abstract
We present QLoRA, an efficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance. QLoRA backpropagates gradients through a frozen, 4-bit quantized pretrained language model into Low Rank Adapters~(LoRA). Our best model family, which we name Guanaco, outperforms all previous openly released models on the Vicuna benchmark, reaching 99.3% of the performance level of ChatGPT while only requiring 24 hours of finetuning on a single GPU. QLoRA introduces a number of innovations to save memory without sacrificing performance: (a) 4-bit NormalFloat (NF4), a new data type that is information theoretically optimal for normally distributed weights (b) double quantization to reduce the average memory footprint by quantizing the quantization constants, and (c) paged optimziers to manage memory spikes. We use QLoRA to finetune more than 1,000 models, providing a detailed analysis of instruction following and chatbot performance across 8 instruction datasets, multiple model types (LLaMA, T5), and model scales that would be infeasible to run with regular finetuning (e.g. 33B and 65B parameter models). Our results show that QLoRA finetuning on a small high-quality dataset leads to state-of-the-art results, even when using smaller models than the previous SoTA. We provide a detailed analysis of chatbot performance based on both human and GPT-4 evaluations showing that GPT-4 evaluations are a cheap and reasonable alternative to human evaluation. Furthermore, we find that current chatbot benchmarks are not trustworthy to accurately evaluate the performance levels of chatbots. A lemon-picked analysis demonstrates where Guanaco fails compared to ChatGPT. We release all of our models and code, including CUDA kernels for 4-bit training.
Cited by
- Layer-wise LoRA fine-tuning: a similarity metric approach
- \kappa-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating
- On the Convergence of Stochastic Low-Rank Adaptation
- Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
- DeFiScreener: Efficient DeFi Attack Pre-screening in Smart Contracts via Historical Case Matching
- Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts
- Quantize with Confidence? An Empirical Study of Quantization for Code Generation
- Stabilizing Native Low-Rank LLM Pretraining
- A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents
- Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
- DINS-IO: Learned Inertial Odometry via Differentiable INS Consistency
- The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability
- ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
- The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation
- Semantic-Aware Data-Aided Channel Estimation with Large Language Models for MIMO Systems
- Chebyshev Manifold Adaptation
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA
- Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
- CommitLLM: A Fine-Tuned Pipeline for Git Commit Message Generation
- What Transfers Under Source Shift? Definitions, Examples, and Fine-Tuning for Climate Disclosure Classification
- OR Else: A Differentiable Trust Region for Policy Optimization
- Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
- Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
- Automating structural reliability analysis with a multi-agent large language model framework
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- Transferable Low-Rank Convolutional Bases for Onboarding Unseen Medical Imaging Modalities
- Long-Context Fine-Tuning with Limited VRAM
- Posts of Peril: Detecting Information About Hazards in Text
- One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
- QDA-SQL: Questions Enhanced Dialogue Augmentation for Multi-Turn Text-to-SQL
- Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
- Program-as-Weights: A Programming Paradigm for Fuzzy Functions
- Large Language Model-Enhanced Multi-hop Parallel Image Semantic Communication
- FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
- Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures
- Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
- Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
- In-Context Probing for Membership Inference in Fine-Tuned Language Models
- The 4/δ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
- Measuring Scalar Constructs in Social Science with LLMs
- A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
- Small Language Models are the Future of Agentic AI
- 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
- LLM Social Simulations Are a Promising Research Method
- Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
- Training Large Neural Networks With Low-Dimensional Error Feedback
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
- The Appeal and Reality of Recycling LoRAs with Adaptive Merging
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- Reservoir Computing inspired Matrix Multiplication-free Language Model
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- Hierarchical Pedagogical Oversight: A Multi-Agent Adversarial Framework for Reliable AI Tutoring
- AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
- Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
- Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh
- TARGET: Automated Scenario Generation from Traffic Rules for Testing Autonomous Vehicles via Validated LLM-Guided Knowledge Extraction
- The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning
- Strategy-Aware Parameter-Efficient Adaptation for LLM-based Auto-Bidding
- AlloBench: Measuring Online Tool Allocation Capability in LLM Agents
- Wrong Design Intent Is Worse Than None: A Derangement-Control Diagnosis of Header Conditioning in CAD Program Completion
- Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age
- Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
- AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
- Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation
- Retrieval-based and Fine-tuned LLM Approaches for Industrial Asset Health Monitoring and Decision Support
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
- Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting
- Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
- LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
- Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
- RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
- Making Large Language Models Efficient Dense Retrievers
- Can abstract concepts from LLM improve SLM performance?
- Generation of Programmatic Rules for Document Forgery Detection Using Large Language Models
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- Code2Doc: A Quality-First Curated Dataset for Code Documentation
- From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- On the Convergence Rate of LoRA Gradient Descent
- Shuttling Compiler for Trapped-Ion Quantum Computers Based on Large Language Models
- Parameter-Efficient Fine-Tuning for HAR: Integrating LoRA and QLoRA into Transformer Models
- Key-Conditioned Orthonormal Transform Gating (K-OTG): Multi-Key Access Control with Hidden-State Scrambling for LoRA-Tuned Models
- Governance-Aware Hybrid Fine-Tuning for Multilingual Large Language Models
- RecipeMasterLLM: Revisiting RoboEarth in the Era of Large Language Models
- Subjective Question Generation and Answer Evaluation using NLP
- Bridging Training and Merging Through Momentum-Aware Optimization
- Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL
- From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
- Per-Axis Weight Deltas for Frequent Model Updates
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- Georeferencing complex relative locality descriptions with large language models
- Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
- NL2SpaTiaL: Generating Geometric Spatio-Temporal Logic Specifications from Natural Language for Manipulation Tasks
- Temporal Tokenization Strategies for Event Sequence Modeling with Large Language Models
- Advancing Bangla Machine Translation Through Informal Datasets
- Socratic Students: Teaching Language Models to Learn by Asking Questions
- Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- Instruction-Tuning Open-Weight Language Models for BPMN Model Generation
- REMODEL-LLM: Transforming C code to Java using LLMs
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Towards Accessible Physical AI: LoRA-Based Fine-Tuning of VLA Models for Real-World Robot Control
- Knowledge Graph Enrichment and Reasoning for Nobel Laureates
- Beyond Real Weights: Hypercomplex Representations for Stable Quantization
- Secure or Suspect? Investigating Package Hallucinations of Shell Command in Original and Quantized LLMs
- LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
- Generalized Referring Expression Segmentation on Aerial Photos
- Large Language Model-Based Generation of Discharge Summaries
- TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- Optimizing LLMs Using Quantization for Mobile Execution
- The Road of Adaptive AI for Precision in Cybersecurity
- Reflection Removal through Efficient Adaptation of Diffusion Transformers
- ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications
- ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- LPCD: Unified Framework from Layer-Wise to Submodule Quantization
- InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages
- Unsupervised decoding of encoded reasoning using language model interpretability
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Responsible LLM Deployment for High-Stake Decisions by Decentralized Technologies and Human-AI Interactions
- Minimal-Edit Instruction Tuning for Low-Resource Indic GEC
- Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
- Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- Gender Bias in Emotion Recognition by Large Language Models
- EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
- Time Travel: LLM-Assisted Semantic Behavior Localization with Git Bisect
- Understanding Task Transfer in Vision-Language Models
- DELTA: Language Diffusion-based EEG-to-Text Architecture
- A Systematic Study of Compression Ordering for Large Language Models
- RFX: High-Performance Random Forests with GPU Acceleration and QLORA Compression
- Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- Principled Design of Interpretable Automated Scoring for Large-Scale Educational Assessments
- Empa: An AI-Powered Virtual Mentor for Developing Global Collaboration Skills in HPC Education
- EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
- The Impact of Quantization on Large Reasoning Model Reinforcement Learning
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- MCAQ-YOLO: Morphological Complexity-Aware Quantization for Efficient Object Detection with Curriculum Learning
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- Privacy Preserving Ordinal-Meta Learning with VLMs for Fine-Grained Fruit Quality Prediction
- Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
- GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
- Reinforcing Trustworthiness in Multimodal Emotional Support Systems
- Community-Aligned Behavior Under Uncertainty: Evidence of Epistemic Stance Transfer in LLMs
- Advanced Black-Box Tuning of Large Language Models with Limited API Calls
- SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification
- Bench360: Benchmarking Local LLM Inference from 360°
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- Hierarchical structure understanding in complex tables with VLLMs: a benchmark and experiments
- SRE-Llama -- Fine-Tuned Meta's Llama LLM, Federated Learning, Blockchain and NFT Enabled Site Reliability Engineering(SRE) Platform for Communication and Networking Software Services
- National Institute on Aging PREPARE Challenge: Early Detection of Cognitive Impairment Using Speech -- The SpeechCARE Solution
- Designing Beyond Language: Sociotechnical Barriers in AI Health Technologies for Limited English Proficiency
- Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
- TuckA: Hierarchical Compact Tensor Experts for Efficient Fine-Tuning
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- Random Initialization of Gated Sparse Adapters
- Continual Learning, Not Training: Online Adaptation For Agents
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers
- Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement
- LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
- Predicate Renaming via Large Language Models
- AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
- NeuronMM: High-Performance Matrix Multiplication for LLM Inference on AWS Trainium
- Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
- Benchmarking Generative AI Against Bayesian Optimization for Constrained Multi-Objective Inverse Design
- How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model
- MemSFT: Mitigating Alignment Tax with an External Parametric Memory
- Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
- PowerScale: Energy-Efficient Geo-Distributed Model Training with Federated Datacenter Power
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Torus embeddings
- Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
- JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
- ReviewSense: Transforming Customer Review Dynamics into Actionable Business Insights
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFT
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- MASPRM: Multi-Agent System Process Reward Model
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning
- FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
- BaZi-Based Character Simulation Benchmark: Evaluating AI on Temporal and Persona Reasoning
- Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks
- Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
- SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning
- Text to Trust: Evaluating Fine-Tuning and LoRA Trade-offs in Language Models for Unfair Terms of Service Detection
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
- Low-Precision Streaming PCA
- Performance Trade-offs of Optimizing Small Language Models for E-Commerce
- α-LoRA: Effective Fine-Tuning via Base Model Rescaling
- Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models
- REx86: A Local Large Language Model for Assisting in x86 Assembly Reverse Engineering
- Hierarchical Sequence Iteration for Heterogeneous Question Answering
- SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- An Empirical Study of Sample Selection Strategies for Large Language Model Repair
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- CoSense-LLM: Semantics at the Edge with Cost- and Uncertainty-Aware Cloud-Edge Cooperation
- Latent Space Factorization in LoRA
- ELUTQ: Efficient LUT-Aware Quantization for Deploying Large Language Models on Edge Devices
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models
- CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- Embodied Navigation with Auxiliary Task of Action Description Prediction
- Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
- Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
- NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
- Model Context Contracts - MCP-Enabled Framework to Integrate LLMs With Blockchain Smart Contracts
- TaxoAlign: Scholarly Taxonomy Generation Using Language Models
- ParaVul: A Parallel Large Language Model and Retrieval-Augmented Framework for Smart Contract Vulnerability Detection
- Qomhra: A Bilingual Irish and English Large Language Model
- Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Finetuning LLMs for EvaCun 2025 token prediction shared task
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- Directional Reasoning Injection for Fine-Tuning MLLMs
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
- A Matter of Representation: Towards Graph-Based Abstract Code Generation
- RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
- CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
- CoRA: Covariate-Aware Adaptation of Time Series Foundation Models
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
- EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
- QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
- FactAppeal: Identifying Epistemic Factual Appeals in News Media
- ADiP: Adaptive Precision Systolic Array for Matrix Multiplication Acceleration
- Traj-CoA: Patient Trajectory Modeling via Chain-of-Agents for Lung Cancer Risk Prediction
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- CTR-LoRA: Curvature-Aware and Trust-Region Guided Low-Rank Adaptation for Large Language Models
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold
- Randomized Gradient Subspaces for Efficient Large Language Model Training
- True 4-Bit Quantized Convolutional Neural Network Training on CPU: Achieving Full-Precision Parity
- Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
- Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- Augmenting Dialog with Think-Aloud Utterances for Modeling Individual Personality Traits by LLM
- Large Language Models Do NOT Really Know What They Don't Know
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- A Unified Biomedical Named Entity Recognition Framework with Large Language Models
- Active Model Selection for Large Language Models
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Interleaved Learning and Exploration: A Self-Adaptive Fuzz Testing Framework for MLIR
- NurseLLM: The First Specialized Language Model for Nursing
- A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
- MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
- ConstraintLLM: A Neuro-Symbolic Framework for Industrial-Level Constraint Programming
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- AI-to-AI Feedback: Amplified Intelligence — Prior Art and Governance Implications for Multi-Model Advisory Architectures
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRA
- FocusMed: A Large Language Model-based Framework for Enhancing Medical Question Summarization with Focus Identification
- Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
- Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
- Referring Expression Comprehension for Small Objects
- Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text
- Memory-Efficient Backpropagation for Fine-Tuning LLMs on Resource-Constrained Mobile Devices
- MALF: A Multi-Agent LLM Framework for Intelligent Fuzzing of Industrial Control Protocols
- Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
- POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
- Integrating AI and Ensemble Forecasting: Explainable Materials Planning with Scorecards and Trend Insights for a Large-Scale Manufacturer
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model
- Efficient Layer-wise LLM Fine-tuning for Revision Intention Prediction
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
- Communication-Efficient and Accurate Approach for Aggregation in Federated Low-Rank Adaptation
- Explaining novel senses using definition generation with open language models
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- Fair Classification by Direct Intervention on Operating Characteristics
- Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs
- Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
- Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- Edge-FIT: Federated Instruction Tuning of Quantized LLMs for Privacy-Preserving Smart Home Environments
- Tequila: Trapping-free Ternary Quantization for Large Language Models
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- VeriGRAG: Enhancing LLM-Based Verilog Code Generation with Structure-Aware Soft Prompts
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
- Effective Quantization of Muon Optimizer States
- Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
- SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
- Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Enhancing Low-Rank Adaptation with Structured Nonlinear Transformations
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
- DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
- Large AI Model-Enabled Generative Semantic Communications for Image Transmission
- Large Language Models for Real-World IoT Device Identification
- CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
- Memory Efficient Tabular Foundation Models
- Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
- Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
- GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- GSTM-HMU: Generative Spatio-Temporal Modeling for Human Mobility Understanding
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- Model selection meets clinical semantics: Optimizing ICD-10-CM prediction via LLM-as-Judge evaluation, redundancy-aware sampling, and section-aware fine-tuning
- Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- Learning to vary: Teaching LMs to reproduce human linguistic variability in next-word prediction
- QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
- SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
- ACCeLLiuM: Supervised Fine-Tuning for Automated OpenACC Pragma Generation
- Can an Individual Manipulate the Collective Decisions of Multi-Agents?
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes
- Semantic Representation Attack against Aligned Large Language Models
- From Hype to Insight: Rethinking Large Language Model Integration in Visual Speech Recognition
- LLM Jailbreak Detection for (Almost) Free!
- Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
- An LLM-based multi-agent framework for agile effort estimation
- AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
- A Multi-Component AI Framework for Computational Psychology: From Robust Predictive Modeling to Deployed Generative Dialogue
- Multi-Model Synthetic Training for Mission-Critical Small Language Models
- Routing Distilled Knowledge via Mixture of LoRA Experts for Large Language Model based Bundle Generation
- A Systematic Evaluation of Parameter-Efficient Fine-Tuning Methods for the Security of Code LLMs
- Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning
- Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
- Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web Agents
- Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges
- UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
- PHLoRA: data-free Post-hoc Low-Rank Adapter extraction from full-rank checkpoint
- CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
- RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment
- Learning from Diverse Reasoning Paths with Routing and Collaboration
- Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- ALIGNS: Unlocking nomological networks in psychological measurement through a large language model
- MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communities
- Interpreting the Effects of Quantization on LLMs
- Let's Roleplay: Examining LLM Alignment in Collaborative Dialogues
- Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
- X-SQL: Expert Schema Linking and Understanding of Text-to-SQL with Multi-LLMs
- Using Contrastive Learning to Improve Two-Way Reasoning in Large Language Models: The Obfuscation Task as a Case Study
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
- Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
- MM-ARC: Multimodal Adaptive Routing of Capital with Robustness-Audited Strategy Pools
- MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
- Holographic Knowledge Manifolds: A Novel Pipeline for Continual Learning Without Catastrophic Forgetting in Large Language Models
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models
- From Injection to Defense: Constructing Edit-Based Fingerprints for Large Language Models
- Binary Quantization For LLMs Through Dynamic Grouping
- Gradient Estimation Methods of Approximate Multipliers for High-Accuracy Retraining of Deep Learning Models
- RoboBuddy in the Classroom: Exploring LLM-Powered Social Robots for Storytelling in Learning and Integration Activities
- Generative AI for Crystal Structures: A Review
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Cloud-Device Collaborative Agents for Sequential Recommendation
- Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
- Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- Language-Aware Information Maximization for Transductive Few-Shot CLIP
- Do Cognitively Interpretable Reasoning Traces Improve LLM Performance?
- QR-LoRA: QR-Based Low-Rank Adaptation for Efficient Fine-Tuning of Large Language Models
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- Integrating Large Language Models with Network Optimization for Interactive and Explainable Supply Chain Planning: A Real-World Case Study
- The Uneven Impact of Post-Training Quantization in Machine Translation
- LLM Chatbot-Creation Approaches
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Research Challenges in Relational Database Management Systems for LLM Queries
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
- LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech
- Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
- APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
- Enhancing Model Privacy in Federated Learning with Random Masking and Quantization
- DrugReasoner: Interpretable Drug Approval Prediction with a Reasoning-augmented Language Model
- End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
- HebID: Detecting Social Identities in Hebrew-language Political Text
- QU-NLP at QIAS 2025 Shared Task: A Two-Phase LLM Fine-Tuning and Retrieval-Augmented Generation Approach for Islamic Inheritance Reasoning
- RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget
- Calibrated Semantic Diffusion: A p-Laplacian Synthesis with Learnable Dissipation, Quantified Constants, and Graph-Aware Calibration
- ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
- Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
- Breaking Language Barriers: Equitable Performance in Multilingual Language Models
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment
- MAD: A Benchmark for Multi-Turn Audio Dialogue Fact-Checking
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering
- Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
- MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
- Large Language Models for Summarizing Czech Historical Documents and Beyond
- Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Data-Driven Discovery of Interpretable Kalman Filter Variants through Large Language Models and Genetic Programming
- Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
- LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
- NEFMind: Parameter-Efficient Fine-Tuning of Open-Source LLMs for Telecom APIs Automation
- UWB at WASSA-2024 Shared Task 2: Cross-lingual Emotion Detection
- LLaMA-Based Models for Aspect-Based Sentiment Analysis
- Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
- Learning User Preferences for Image Generation Model
- Large Language Models for Subjective Language Understanding: A Survey
- Large Language Models for Czech Aspect-Based Sentiment Analysis
- LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
- Surgical Knowledge Rewrite in Compact LLMs: An 'Unlearn-then-Learn' Strategy with (IA3) for Localized Factual Modulation and Catastrophic Forgetting Mitigation
- The Trauma THOMPSON Dataset for Real-World Emergency AI
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- Multimodal Fact Checking with Unified Visual, Textual, and Contextual Representations
- Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
- HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
- FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
- Modelling and Classifying the Components of a Literature Review
- Fine-tuning for Better Few Shot Prompting: An Empirical Comparison for Short Answer Grading
- Sotopia-RL: Reward Design for Social Intelligence
- EmbedGrad: Gradient-Based Prompt Optimization in Embedding Space for Large Language Models
- Key-Augmented Neural Triggers for Knowledge Sharing
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
- MoKA: Mixture of Kronecker Adapters
- PLoRA: Efficient LoRA Hyperparameter Tuning for Large Models
- LOST: Low-rank and Sparse Pre-training for Large Language Models
- Traffic-R1: Reinforced LLMs Bring Human-Like Reasoning to Traffic Signal Control Systems
- Everyone Contributes! Incentivizing Strategic Cooperation in Multi-LLM Systems via Sequential Public Goods Games
- Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Prompting Large Language Models to Detect Dementia Family Caregivers
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- Convergence Analysis of Aggregation-Broadcast in LoRA-enabled Distributed Fine-Tuning
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
- How Quantization Impacts Privacy Risk on LLMs for Code?
- Enabling Few-Shot Alzheimer's Disease Diagnosis on Biomarker Data with Tabular LLMs
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- Automated Label Placement on Maps via Large Language Models
- Compression Strategies for Efficient Multimodal LLMs in Medical Contexts
- EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
- Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour
- LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
- Infogen: Generating Complex Statistical Infographics from Documents
- Ari Holtzman [wikipedia]
- Post-training of large language models [wikipedia]
Discussions
Related