AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
2023/06/01 by Ji Lin, Lin, Ji, Jiaming Tang +18 · 1 voice · 298 citations
Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2306.00978
Abstract
Large language models (LLMs) have transformed numerous AI applications. On-device LLM is becoming increasingly important: running LLMs locally on edge devices can reduce the cloud computing cost and protect users' privacy. However, the astronomical model size and the limited hardware resource pose significant deployment challenges. We propose Activation-aware Weight Quantization (AWQ), a hardware-friendly approach for LLM low-bit weight-only quantization. AWQ finds that not all weights in an LLM are equally important. Protecting only 1% salient weights can greatly reduce quantization error. To identify salient weight channels, we should refer to the activation distribution, not weights. To avoid the hardware-inefficient mix-precision quantization, we mathematically derive that scaling up the salient channels can reduce the quantization error. AWQ employs an equivalent transformation to scale the salient weight channels to protect them. The scale is determined by collecting the activation statistics offline. AWQ does not rely on any backpropagation or reconstruction, so it generalizes to different domains and modalities without overfitting the calibration set. AWQ outperforms existing work on various language modeling and domain-specific benchmarks (coding and math). Thanks to better generalization, it achieves excellent quantization performance for instruction-tuned LMs and, for the first time, multi-modal LMs. Alongside AWQ, we implement TinyChat, an efficient and flexible inference framework tailored for 4-bit on-device LLM/VLMs. With kernel fusion and platform-aware weight packing, TinyChat offers more than 3x speedup over the Huggingface FP16 implementation on both desktop and mobile GPUs. It also democratizes the deployment of the 70B Llama-2 model on mobile GPUs.
Cited by
- Active Perception Agent for Omnimodal Audio-Video Understanding
- Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
- Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
- Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
- MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
- CausalGate: Causal Importance Distillation for Transformer Module Pruning
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
- MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models
- GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
- TimeBill: Time-Budgeted Inference for Large Language Models
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- Towards Efficient Agents: A Co-Design of Inference Architecture and System
- Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation
- CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
- Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
- Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
- CKA-Guided Modular Quantization: Beyond Bit-Width to Algorithmic Diversity
- Dynamic Rebatching for Efficient Early-Exit Inference with DREX
- SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
- Arithmetic-Intensity-Aware Quantization
- SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
- Low-Rank Compression of Language Models via Differentiable Rank Selection
- Compressed Causal Reasoning: Quantization and GraphRAG Effects on Interventional and Counterfactual Accuracy
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
- Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning
- ELANA: A Simple Energy and Latency Analyzer for LLMs
- KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
- Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
- SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
- Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
- ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Data-Free Pruning of Self-Attention Layers in LLMs
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- Understanding and Harnessing Sparsity in Unified Multimodal Models
- FOVA: Offline Federated Reinforcement Learning with Mixed-Quality Data
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
- A Systematic Characterization of LLM Inference on GPUs
- LPCD: Unified Framework from Layer-Wise to Submodule Quantization
- Intrinsic Structure as a Proxy for Saliency: SVD-Based Weight Preservation for Mixed-Precision Quantization in Large Language Models
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
- Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- Efficient Kernel Mapping and Comprehensive System Evaluation of LLM Acceleration on a CGLA
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- Experts are all you need: A Composable Framework for Large Language Model Inference
- Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
- The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
- Artwork Interpretation with Vision Language Models: A Case Study on Emotions and Emotion Symbols
- Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
- VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
- Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
- Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Arctic-Extract Technical Report
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
- The Impact of Quantization on Large Reasoning Model Reinforcement Learning
- Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone
- T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
- OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs
- BitSnap: Checkpoint Sparsification and Quantization in LLM Training
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
- Bench360: Benchmarking Local LLM Inference from 360°
- Alignment-Aware Quantization for LLM Safety
- SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
- P3-LLM: An Integrated NPU-PIM Accelerator for LLM Inference Using Hybrid Numerical Formats
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
- QuAnTS: Question Answering on Time Series
- Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
- Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- Reasoning Planning for Language Models
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?
- LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
- INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- RDQ: Residual Distribution Quantization for Large Language Models
- LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
- HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
- Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
- Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- Pie: A Programmable Serving System for Emerging LLM Applications
- PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
- FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
- BitSkip: An Empirical Analysis of Quantization and Early Exit Composition
- PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Performance Trade-offs of Optimizing Small Language Models for E-Commerce
- Correlation Dimension of Auto-Regressive Large Language Models
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- GRATING: Low-Latency and Memory-Efficient Semantic Selection on Device
- SpikeFit: Towards Optimal Deployment of Spiking Networks on Neuromorphic Hardware
- HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
- CoSense-LLM: Semantics at the Edge with Cost- and Uncertainty-Aware Cloud-Edge Cooperation
- Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
- ELUTQ: Efficient LUT-Aware Quantization for Deploying Large Language Models on Edge Devices
- CPSVD: Enhancing Large Language Model Compression via Column-Preserving Singular Value Decomposition
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval
- From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- Neuronal Group Communication for Efficient Neural representation
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- FraQAT: Quantization Aware Training with Fractional bits
- AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- BitNet Distillation
- Optimizing Storage Overhead of User Behavior Log for ML-embedded Mobile Apps
- CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
- Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
- QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
- FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management
- XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
- Neural Weight Compression for Language Models
- Automating Structural Engineering Workflows with Large Language Model Agents
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- SASER: Stego attacks on open-source LLMs
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model
- Accelerating Attention with Basis Decomposition
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
- Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Training Dynamics Impact Post-Training Quantization Robustness
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
- To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Accelerating LLM Inference with Precomputed Query Storage
- SAIL: SRAM-Accelerated LLM Inference System with Lookup-Table-based GEMV
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- Layer-wise dynamic rank for compressing large language models
- Fingerprinting LLMs via Prompt Injection
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
- Tequila: Trapping-free Ternary Quantization for Large Language Models
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- Effective Quantization of Muon Optimizer States
- PT2-LLM: Post-Training Ternarization for Large Language Models
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
- SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
- InfiR2: A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models
- AxLLM: accelerator architecture for large language models with computation reuse capability
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- Compute-Optimal Quantization-Aware Training
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
- Quantized Visual Geometry Grounded Transformer
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
- LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- Memory Efficient Tabular Foundation Models
- Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
- Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
- WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
- QuantWAMs: Calibrating at the Right Granularity for World Action Models
- Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
- Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning
- Linguistic properties and model scale in brain encoding: from small to compressed language models
- Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering
- Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
- SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
- CLMTracing: Black-box User-level Watermarking for Code Language Model Tracing
- SBVR: Summation of BitVector Representation for Efficient LLM Quantization
- CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
- Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
- Exploring and Reshaping the Weight Distribution in LLM
- Spatio-Temporal Pruning for Compressed Spiking Large Language Models
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- Compass-v3: Scaling Domain-Specific LLMs for Multilingual E-Commerce in Southeast Asia
- Building Large-Scale English-Romanian Literary Translation Resources with Open Models
- Interpreting the Effects of Quantization on LLMs
- DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
- COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
- FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
- LoaQ: Layer-wise Output Approximation Quantization
- AttestLLM: Efficient Attestation Framework for Billion-scale On-device LLMs
- A Node-Aware Dynamic Quantization Approach for Graph Collaborative Filtering
- Delta Activations: A Representation for Finetuned Large Language Models
- Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
- QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception
- Binary Quantization For LLMs Through Dynamic Grouping
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- Dynamic Sparse Attention on Mobile SoCs
- LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- The Uneven Impact of Post-Training Quantization in Machine Translation
- cMALC-D: Contextual Multi-Agent LLM-Guided Curriculum Learning with Diversity-Based Context Blending
- Beacon: Post-Training Quantization with Integrated Grid Selection
- LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
- Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Scaling Laws for Task-Stratified Knowledge in Post-Training Quantized Large Language Models
- Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance
- How Quantization Shapes Bias in Large Language Models
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- Quantum Relational Knowledge Distillation
- LLM Compression: How Far Can We Go in Balancing Size and Performance?
- CSGO: Generalized Optimization for Cold Start in Wireless Collaborative Edge LLM Systems
- SoK: Data Minimization in Machine Learning
- SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
- Leveraging OS-Level Primitives for Robotic Action Management
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- P/D-Device: Disaggregated Large Language Model between Cloud and Devices
- OverFill: Two-Stage Models for Efficient Language Model Decoding
- Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
- MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
- Pushing the Envelope of LLM Inference on AI-PC
- DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
- Optimal Brain Connection: Towards Efficient Structural Pruning
- Camel: Energy-Aware LLM Inference on Resource-Constrained Devices
- Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
- Fine-Tuning Vision-Language Models for Markdown Conversion of Financial Tables in Malaysian Audited Financial Reports
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- Systematic Evaluation of Optimization Techniques for Long-Context Language Models
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- KLLM: Fast LLM Inference with K-Means Quantization
- Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
- Compression Strategies for Efficient Multimodal LLMs in Medical Contexts
Discussions
Related