OPT: Open Pre-trained Transformer Language Models
2022/05/02 by Susan Zhang, Stephen Roller, Zhang, Susan +35 · 2 voices · 258 citations
#cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2205.01068
Abstract
Large language models, which are often trained for hundreds of thousands of compute days, have shown remarkable capabilities for zero- and few-shot learning. Given their computational cost, these models are difficult to replicate without significant capital. For the few that are available through APIs, no access is granted to the full model weights, making them difficult to study. We present Open Pre-trained Transformers (OPT), a suite of decoder-only pre-trained transformers ranging from 125M to 175B parameters, which we aim to fully and responsibly share with interested researchers. We show that OPT-175B is comparable to GPT-3, while requiring only 1/7th the carbon footprint to develop. We are also releasing our logbook detailing the infrastructure challenges we faced, along with code for experimenting with all of the released models.
Cited by
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes
- Test Case Prioritization for DNNs via Neural Collapse Instability
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- Statistical Inference for Generative Model Comparison
- Abstraction Induces the Brain Alignment of Language and Speech Models
- (A)iSpy: Parasitic Trojans for Machine Learning Infrastructure
- DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
- BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
- Perturbation is All You Need for Extrapolating Language Models
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- Interactive Training 2: Auditable Control Plane for Live Model Training
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
- Sparser, Faster, Lighter Transformer Language Models
- A Survey on Diffusion Language Models
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
- The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement
- LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
- From Fake Focus to Real Precision: Confusion-Driven Adversarial Attention Learning in Transformers
- DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
- Epistemic diversity across language models mitigates knowledge collapse
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- Lyra: A Hardware-Accelerated RISC-V Verification Framework with Generative Model-Based Processor Fuzzing
- Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
- CurvaDion: Curvature-Adaptive Distributed Orthonormalization
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
- Bandwidth-Aware Network Topology Optimization for Decentralized Learning
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- Do Generalisation Results Generalise?
- BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Large Language Models as Generalist Policies for Network Optimization
- TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
- Experts are all you need: A Composable Framework for Large Language Model Inference
- Towards Audio Token Compression in Large Audio Language Models
- LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
- CDLM: Consistency Diffusion Language Models For Faster Sampling
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
- A cross-species neural foundation model for end-to-end speech decoding
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
- Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
- GPS: General Per-Sample Prompter
- 10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
- BitSnap: Checkpoint Sparsification and Quantization in LLM Training
- OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
- Dynamic Temperature Scheduler for Knowledge Distillation
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
- LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
- ProcGen3D: Learning Neural Procedural Graph Representations for Image-to-3D Reconstruction
- GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
- HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- The Future of Fully Homomorphic Encryption System: from a Storage I/O Perspective
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- UMDAM: A Unified Data Layout and DRAM Address Mapping for Heterogenous NPU-PIM
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
- A CPU-Centric Perspective on Agentic AI
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
- Efficient semantic uncertainty quantification in language models via diversity-steered sampling
- Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing
- TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
- Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
- Relative-Based Scaling Law for Neural Language Models
- Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
- What is the Best Sequence Length for BABYLM?
- On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
- Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
- Learning Human-Object Interaction as Groups
- BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining
- Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Graph4MM: Weaving Multimodal Learning with Structural Information
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Zeroth-Order Sharpness-Aware Learning with Exponential Tilting
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
- Towards Reversible Model Merging For Low-rank Weights
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- Readout Representation: Redefining Neural Codes by Input Recovery
- An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
- Bolster Hallucination Detection via Prompt-Guided Data Augmentation
- Large Language Model-Empowered Channel Prediction and Predictive Beamforming for LEO Satellite Communications
- Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- Softmax ≥ Linear: Transformers may learn to classify in-context by kernel gradient descent
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- On the Provable Performance Guarantee of Efficient Reasoning Models
- Detecting Post-generation Edits to Watermarked LLM Outputs via Combinatorial Watermarking
- Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence
- Cocoon: A System Architecture for Differentially Private Training with Correlated Noises
- Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations
- Mid-Training of Large Language Models: A Survey
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Mixture of Neuron Experts
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- LongTail-Swap: benchmarking language models' abilities on rare words
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Neural Correlates of Language Models Are Specific to Human Language
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
- HiSpec: Hierarchical Speculative Decoding for LLMs
- Scaling Spoken Language Models with Syllabic Speech Tokenization
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- SAIL: SRAM-Accelerated LLM Inference System with Lookup-Table-based GEMV
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- Negative Pre-activations Differentiate Syntax
- Tequila: Trapping-free Ternary Quantization for Large Language Models
- GeoBS: Information-Theoretic Quantification of Geographic Bias in AI Models
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- Knowledge distillation through geometry-aware representational alignment
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- PT2-LLM: Post-Training Ternarization for Large Language Models
- LLM Watermark Evasion via Bias Inversion
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- LIMI: Less is More for Agency
- SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
- CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
- Character-Level Perturbations Disrupt LLM Watermarks
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- Interpreting the Effects of Quantization on LLMs
- SMooGPT: Stylized Motion Generation using Large Language Models
- RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- Behavioral Fingerprinting of Large Language Models
- Dynamic Sparse Attention on Mobile SoCs
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generation
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Subjective Behaviors and Preferences in LLM: Language of Browsing
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Discrete Optimization of Min-Max Violation and its Applications Across Computational Sciences
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
- Puppeteer: Rig and Animate Your 3D Models
- ORBIT: An Object Property Reasoning Benchmark for Visual Inference Tasks
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- SinLlama -- A Large Language Model for Sinhala
- VertexRegen: Mesh Generation with Continuous Level of Detail
- Semantic-Enhanced Time-Series Forecasting via Large Language Models
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
- Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning
- Approaching the integration of large language models in the parliamentary workspace
- When a Paper Has 1000 Authors: Rethinking Citation Metrics in the Era of LLMs
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- Understanding the Landscape of Ampere GPU Memory Errors
- CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- A Bayesian Hybrid Parameter-Efficient Fine-Tuning Method for Large Language Models
- KLLM: Fast LLM Inference with K-Means Quantization
- When Truthful Representations Flip Under Deceptive Instructions?
- Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
- Shapley Uncertainty in Natural Language Generation
- Do Large Language Models Understand Morality Across Cultures?
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- List of large language models [wikipedia]
Discussions
Related