OPT: Open Pre-trained Transformer Language Models
2022/05/02 by Susan Zhang, Stephen Roller, Zhang, Susan +35 · 2 voices · 519 citations
Computer Science · #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2205.01068
arxiv created 2022/06/21 · arxiv updated 2022/06/22
Abstract
Large language models, which are often trained for hundreds of thousands of compute days, have shown remarkable capabilities for zero- and few-shot learning. Given their computational cost, these models are difficult to replicate without significant capital. For the few that are available through APIs, no access is granted to the full model weights, making them difficult to study. We present Open Pre-trained Transformers (OPT), a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters, which we aim to fully and responsibly share with interested researchers. We show that OPT-175B is comparable to GPT-3, while requiring only 1/7th the carbon footprint to develop. We are also releasing our logbook detailing the infrastructure challenges we faced, along with code for experimenting with all of the released models.
Cited by
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes
- Test Case Prioritization for DNNs via Neural Collapse Instability
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- Statistical Inference for Generative Model Comparison
- Abstraction Induces the Brain Alignment of Language and Speech Models
- (A)iSpy: Parasitic Trojans for Machine Learning Infrastructure
- DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
- BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
- Perturbation is All You Need for Extrapolating Language Models
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- Interactive Training 2: Auditable Control Plane for Live Model Training
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
- Sparser, Faster, Lighter Transformer Language Models
- A Survey on Diffusion Language Models
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
- The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement
- LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
- Calibrating Transformer Attention via Task-Space Sensitivity Feedback
- DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
- Epistemic diversity across language models mitigates knowledge collapse
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- Lyra: A Hardware-Accelerated RISC-V Verification Framework with Generative Model-Based Processor Fuzzing
- Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
- CurvaDion: Curvature-Adaptive Distributed Orthonormalization
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
- Bandwidth-Aware Network Topology Optimization for Decentralized Learning
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- Do Generalisation Results Generalise?
- BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Large Language Models as Generalist Policies for Network Optimization
- TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
- Experts are all you need: A Composable Framework for Large Language Model Inference
- Towards Audio Token Compression in Large Audio Language Models
- LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
- CDLM: Consistency Diffusion Language Models For Faster Sampling
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
- A cross-species neural foundation model for end-to-end speech decoding
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
- Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
- GPS: General Per-Sample Prompter
- 10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
- BitSnap: Checkpoint Sparsification and Quantization in LLM Training
- OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
- Dynamic Temperature Scheduler for Knowledge Distillation
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
- LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
- ProcGen3D: Learning Neural Procedural Graph Representations for Image-to-3D Reconstruction
- GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
- HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- The Future of Fully Homomorphic Encryption System: from a Storage I/O Perspective
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
- Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
- Efficient semantic uncertainty quantification in language models via diversity-steered sampling
- Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing
- TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
- Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
- Relative-Based Scaling Law for Neural Language Models
- Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
- What is the Best Sequence Length for BABYLM?
- On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
- Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
- Learning Human-Object Interaction as Groups
- BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
- Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Graph4MM: Weaving Multimodal Learning with Structural Information
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Zeroth-Order Sharpness-Aware Learning with Exponential Tilting
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
- Towards Reversible Model Merging For Low-rank Weights
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- Readout Representation: Redefining Neural Codes by Input Recovery
- An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
- Bolster Hallucination Detection via Prompt-Guided Data Augmentation
- Large Language Model-Empowered Channel Prediction and Predictive Beamforming for LEO Satellite Communications
- Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- Softmax ≥ Linear: Transformers may learn to classify in-context by kernel gradient descent
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- On the Provable Performance Guarantee of Efficient Reasoning Models
- Detecting Post-generation Edits to Watermarked LLM Outputs via Combinatorial Watermarking
- Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence
- Cocoon: A System Architecture for Differentially Private Training with Correlated Noises
- Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations
- Mid-Training of Large Language Models: A Survey
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Mixture of Neuron Experts
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- LongTail-Swap: benchmarking language models' abilities on rare words
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Neural Correlates of Language Models Are Specific to Human Language
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
- HiSpec: Hierarchical Speculative Decoding for LLMs
- Scaling Spoken Language Models with Syllabic Speech Tokenization
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- SAIL: SRAM-Accelerated LLM Inference System with Lookup-Table-based GEMV
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- Negative Pre-activations Differentiate Syntax
- Tequila: Trapping-free Ternary Quantization for Large Language Models
- GeoBS: Information-Theoretic Quantification of Geographic Bias in AI Models
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- Knowledge distillation through geometry-aware representational alignment
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- PT2-LLM: Post-Training Ternarization for Large Language Models
- LLM Watermark Evasion via Bias Inversion
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM Inference
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- LIMI: Less is More for Agency
- SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
- CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
- Character-Level Perturbations Disrupt LLM Watermarks
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- Interpreting the Effects of Quantization on LLMs
- SMooGPT: Stylized Motion Generation using Large Language Models
- RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- Behavioral Fingerprinting of Large Language Models
- ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generation
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Subjective Behaviors and Preferences in LLM: Language of Browsing
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Discrete Optimization of Min-Max Violation and its Applications Across Computational Sciences
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
- Puppeteer: Rig and Animate Your 3D Models
- A Study of Commonsense Reasoning over Visual Object Properties
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- SinLlama -- A Large Language Model for Sinhala
- VertexRegen: Mesh Generation with Continuous Level of Detail
- Semantic-Enhanced Time-Series Forecasting via Large Language Models
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
- Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning
- The NordDRG AI Benchmark for Large Language Models
- Approaching the integration of large language models in the parliamentary workspace
- When a Paper Has 1000 Authors: Rethinking Citation Metrics in the Era of LLMs
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- Understanding the Landscape of Ampere GPU Memory Errors
- CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
- Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- A Bayesian Hybrid Parameter-Efficient Fine-Tuning Method for Large Language Models
- OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration
- When Truthful Representations Flip Under Deceptive Instructions?
- Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Shapley Uncertainty in Natural Language Generation
- Do Large Language Models Understand Morality Across Cultures?
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
- Flora: Effortless Context Construction to Arbitrary Length and Scale
- Modality Agnostic Efficient Long Range Encoder
- SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models
- MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
- PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
- Rethinking Memorization Measures and their Implications in Large Language Models
- Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design
- BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
- BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
- Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
- Detecting LLM-generated Code with Subtle Modification by Adversarial Training
- 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
- PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
- Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage
- SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
- ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
- How Good LLM-Generated Password Policies Are?
- LogTinyLLM: Tiny Large Language Models Based Contextual Log Anomaly Detection
- Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
- AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air
- KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- Post-Training Quantization of Generative and Discriminative LSTM Text Classifiers: A Study of Calibration, Class Balance, and Robustness
- ViSP: A PPO-Driven Framework for Sarcasm Generation with Contrastive Learning
- HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
- SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
- SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
- Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks
- The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
- OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion
- Heterogeneous User Modeling for LLM-based Recommendation
- GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
- Graph Neural Networks as a Substitute for Transformers in Single-Cell Transcriptomics
- Disentangling the Roles of Representation and Selection in Data Pruning
- MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
- DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
- HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
- Time-Masked Transformers with Lightweight Test-Time Adaptation for Neural Speech Decoding
- An AI-native experimental laboratory for autonomous biomolecular engineering
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
- GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
- Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
- Quantize-Sample-and-Verify: LLM Acceleration via Adaptive Edge-Cloud Speculative Decoding
- Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation
- Memory Savings at What Cost? A Study of Alternatives to Backpropagation
- Pay Attention to Small Weights
- LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking
- Characterization and Mitigation of Training Instabilities in Microscaling Formats
- DriveBLIP2: Attention-Guided Explanation Generation for Complex Driving Scenarios
- DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
- Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
- Generalizing vision-language models to novel domains: A comprehensive survey
- From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
- Statistical Multicriteria Evaluation of LLM-Generated Text
- Multi-Amateur Contrastive Decoding for Text Generation
- SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization
- Large Language Models as Psychological Simulators: A Methodological Guide
- Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
- All is Not Lost: LLM Recovery without Checkpoints
- eLLM: Elastic Memory Management Framework for Efficient LLM Serving
- Memory-Efficient Differentially Private Training with Gradient Random Projection
- DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
- Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
- Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
- MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
- Exploring Cultural Variations in Moral Judgments with Large Language Models
- Fed-HeLLo: Efficient Federated Foundation Model Fine-Tuning with Heterogeneous LoRA Allocation
- NoLoCo: No-all-reduce Low Communication Training Method for Large Models
- Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly
- TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
- FREE: Fast and Robust Vision Language Models with Early Exits
- Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction
- Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
- BAQ: Efficient Bit Allocation Quantization for Large Language Models
- Being Strong Progressively! Enhancing Knowledge Distillation of Large Language Models through a Curriculum Learning Framework
- Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
- SoK: Are Watermarks in LLMs Ready for Deployment?
- SECNEURON: Reliable and Flexible Abuse Control in Local LLMs via Hybrid Neuron Encryption
- X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
- Diffusion Model Quantization: A Review
- Kinetics: Rethinking Test-Time Scaling Laws
- Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
- AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
- Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
- Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
- Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
- Parallel CPU-GPU Execution for LLM Inference on Constrained GPUs
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
- STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
- Beyond Text Compression: Evaluating Tokenizers Across Scales
- ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
- Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
- Automatic Calibration for Membership Inference Attack on Large Language Models
- Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples
- SPAP: Structured Pruning via Alternating Optimization and Penalty Methods
- LittleBit: Ultra Low-Bit Quantization via Latent Factorization
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
- Curse of High Dimensionality Issue in Transformer for Long-context Modeling
- ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
- Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
- Beyond path selection: Better LLMs for Scientific Information Extraction with MimicSFT and Relevance and Rule-induced(R2)GRPO
- Highly Efficient and Effective LLMs with Multi-Boolean Architectures
- RISE: Reasoning Enhancement via Iterative Self-Exploration in Multi-hop Question Answering
- Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
- Pretraining Language Models to Ponder in Continuous Space
- Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
- Test-Time Learning for Large Language Models
- MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning
- Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
- ResSVD: Residual Compensated SVD for Large Language Model Compression
- Radio: Rate-Distortion Optimization for Large Language Model Compression
- FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
- An End-to-End Model for Logits-Based Large Language Models Watermarking
- Frictional Agent Alignment Framework: Slow Down and Don't Break Things
- Towards Harmonized Uncertainty Estimation for Large Language Models
- eACGM: Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model
- Rethinking the Understanding Ability across LLMs through Mutual Information
- Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
- Language Model Distillation: A Temporal Difference Imitation Learning Perspective
- μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
- KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
- LatentLLM: Attention-Aware Joint Tensor Compression
- Understanding Gated Neurons in Transformers from Their Input-Output Functionality
- Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
- Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
- Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
- Reinforcement Speculative Decoding for Fast Ranking
- Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
- SELF: Self-Extend the Context Length With Logistic Growth Function
- TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
- Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting
- Incremental Sequence Classification with Temporal Consistency
- LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
- NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
- Revealing Language Model Trajectories via Kullback-Leibler Divergence
- SUS backprop: linear backpropagation algorithm for long inputs in transformers
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association
- Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language Models
- PRL: Prompts from Reinforcement Learning
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion
- Capturing the Effects of Quantization on Trojans in Code LLMs
- Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
- Domain Gating Ensemble Networks for AI-Generated Text Detection
- Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
- Quantum Knowledge Distillation for Large Language Models
- Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
- TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
- GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
- Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
- Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks
- The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
- Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
- Superposition Yields Robust Neural Scaling
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models
- MorphMark: Flexible Adaptive Watermarking for Large Language Models
- Resource-Efficient Language Models: Quantization for Fast and Accessible Inference
- Motif-Mamba: network motif improved mamba for long-range sequence modeling
- Detecting Prefix Bias in LLM-based Reward Models
- Evaluating Financial Sentiment Analysis with Annotators Instruction Assisted Prompting: Enhancing Contextual Interpretation and Stock Prediction Accuracy
- Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
- Scalable LLM Math Reasoning Acceleration with Low-rank Distillation
- Demystifying optimized prompts in language models
- LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps
- MateICL: Mitigating Attention Dispersion in Large-Scale In-Context Learning
- Don't be lazy: CompleteP enables compute-efficient deep transformers
- Fast and Low-Cost Genomic Foundation Models via Outlier Removal
- Learning to Aggregate Zero-Shot LLM Agents for Corporate Disclosure Classification
- CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass
- LLMPrism: Black-box Performance Diagnosis for Production LLM Training Platforms
- Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
- An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
- Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
- UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
- Inside the LLM Word Factory
- Neuron Populations Exhibit Divergent Selectivity with Scale
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
- Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
- SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
- Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
- HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
- AndroidGen: Building an Android Language Agent under Data Scarcity
- Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- The Big Send-off: High Performance Collectives on GPU-based Supercomputers
- Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review
- Leveraging Decoder Architectures for Learned Sparse Retrieval
- Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
- Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
- ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
- SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics
- SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels
- HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
- A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
- Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
- CoheMark: A Novel Sentence-Level Watermark for Enhanced Text Quality
- Language Models Generalize to Human-like Word Order Preferences
- URECA: Unique Region Caption Anything
- On Design Principles for Efficient Heterogeneous DRAM-PIM-GPU Systems
- Your Image Generator Is Your New Private Dataset
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- Hessian of Perplexity for Large Language Models by PyTorch autograd (Open Source)
- Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
- Domain Generalization for Face Anti-spoofing via Content-aware Composite Prompt Engineering
- SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
- Context-Enhanced Contrastive Search for Improved LLM Text Generation
- BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
- W-PCA Based Gradient-Free Proxy for Efficient Search of Lightweight Language Models
- Hardware-based Heterogeneous Memory Management for Large Language Model Inference
- ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
- Analysing the Robustness of Vision-Language-Models to Common Corruptions
- Tinker Tales: Interactive Storytelling Framework for Early Childhood Narrative Development and AI Literacy
- DIDS: Domain Impact-aware Data Sampling for Large Language Model Training
- SLOs-Serve: Optimized Serving of Multi-SLO LLMs
- Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
- Entropy-Guided Watermarking for LLMs: A Test-Time Framework for Robust and Traceable Text Generation
- One Model to Rig Them All: Diverse Skeleton Rigging with UniRig
- A Perplexity and Menger Curvature-Based Approach for Similarity Evaluation of Large Language Models
- SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
- CSPLADE: Learned Sparse Retrieval with Causal Language Models
- Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
- Transferable text data distillation by trajectory matching
- A Tale of Two Learning Algorithms: Multiple Stream Random Walk and Asynchronous Gossip
- DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
- AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
- Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
- HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
- Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
- Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Knowledge Graph-extended Retrieval Augmented Generation for Question Answering
- Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
- Exploring the Effectiveness and Interpretability of Texts in LLM-based Time Series Models
- Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
- List of large language models [wikipedia]
Discussions
Related