Mixed Precision Training
2017/10/10 by Paulius Micikevicius, Sharan Narang, Micikevicius, Paulius +20 · 1 voice · 171 citations
Physics and Astronomy · Computer Science · #Model Reduction and Neural Networks #Numerical Methods and Algorithms #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.1710.03740
Abstract
Deep neural networks have enabled progress in a wide variety of applications. Growing the size of the neural network typically results in improved accuracy. As model sizes grow, the memory and compute requirements for training these models also increases. We introduce a technique to train deep neural networks using half precision floating point numbers. In our technique, weights, activations and gradients are stored in IEEE half-precision format. Half-precision floating numbers have limited numerical range compared to single-precision numbers. We propose two techniques to handle this loss of information. Firstly, we recommend maintaining a single-precision copy of the weights that accumulates the gradients after each optimizer step. This single-precision copy is rounded to half-precision format during training. Secondly, we propose scaling the loss appropriately to handle the loss of information with half-precision gradients. We demonstrate that this approach works for a wide variety of models including convolution neural networks, recurrent neural networks and generative adversarial networks. This technique works for large scale models with more than 100 million parameters trained on large datasets. Using this approach, we can reduce the memory consumption of deep learning models by nearly 2x. In future processors, we can also expect a significant computation speedup using half-precision hardware units.
Citations
Cited by
- Automated Numerical Stability Analysis of Deep Learning Operators
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
- GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
- Economical Jet Taggers -- Equivariant, Slim, and Quantized
- LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- CurvaDion: Curvature-Adaptive Distributed Orthonormalization
- WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
- Generalized Referring Expression Segmentation on Aerial Photos
- Enhanced Chest Disease Classification Using an Improved CheXNet Framework with EfficientNetV2-M and Optimization-Driven Learning
- Curvature-Regularized Variational Autoencoder for 3D Scene Reconstruction from Sparse Depth
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling
- Flexible Gravitational-Wave Parameter Estimation with Transformers
- Learning Eigenstructures of Unstructured Data Manifolds
- Dialect Identification Using Resource-Efficient Fine-Tuning Approaches
- Accelerating Bangla NLP Tasks with Automatic Mixed Precision: Resource-Efficient Training Preserving Model Efficacy
- EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants
- Deep Learning for Restoring MPI System Matrices Using Simulated Training Data
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- SpectraNet: FFT-assisted Deep Learning Classifier for Deepfake Face Detection
- Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
- A Systematic Study of Compression Ordering for Large Language Models
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- MRI Super-Resolution with Deep Learning: A Comprehensive Survey
- Optimizing Federated Learning in the Era of LLMs: Message Quantization and Streaming
- Quant-Trim in Practice: Improved Cross-Platform Low-Bit Deployment on Edge NPUs
- Aligning Generative Music AI with Human Preferences: Methods and Challenges
- Gradient-descent methods for quantum detector tomography
- 10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
- Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
- DPVO-QAT++: Heterogeneous QAT and CUDA Kernel Fusion for High-Performance Deep Patch Visual Odometry
- HyperComplEx: Adaptive Multi-Space Knowledge Graph Embeddings
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- Learning Sparse Label Couplings for Multilabel Chest X-Ray Diagnosis
- DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
- ML-EcoLyzer: Quantifying the Environmental Cost of Machine Learning Inference Across Frameworks and Hardware
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
- Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
- FedSparQ: Adaptive Sparse Quantization with Error Feedback for Robust & Efficient Federated Learning
- Seeing Across Time and Views: Multi-Temporal Cross-View Learning for Robust Video Person Re-Identification
- FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
- HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images
- Training with Fewer Bits: Unlocking Edge LLMs Training with Stochastic Rounding
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- AD-SAM: Fine-Tuning the Segment Anything Vision Foundation Model for Autonomous Driving Perception
- Defeating the Training-Inference Mismatch via FP16
- OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender
- Attentive fine-tuning of Transformers for Translation of low-resourced\n languages @LoResMT 2021
- EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System
- Taking Notes on the Fly Helps BERT Pre-training
- TorchIO: A Python library for efficient loading, preprocessing, augmentation and patch-based sampling of medical images in deep learning
- Conv-Linformer: Boosting Linformer's Performance with Convolution in Small-Scale Settings
- QGAN: Quantized Generative Adversarial Networks
- NeurST: Neural Speech Translation Toolkit
- TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction
- Paris: A Decentralized Trained Open-Weight Diffusion Model
- 8-bit Optimizers via Block-wise Quantization
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Understanding and Improving Fast Adversarial Training
- Revisiting BFloat16 Training
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- Mixed Precision Training of Convolutional Neural Networks using Integer Operations
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- Slalom: Fast, Verifiable and Private Execution of Neural Networks in Trusted Hardware
- Zero-Shot Text-to-Image Generation
- Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- Improving the Straight-Through Estimator with Zeroth-Order Information
- Mixed Precision Training of Neural ODEs
- Selection via Proxy: Efficient Data Selection for Deep Learning
- PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
- Low-Precision Streaming PCA
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- COLD: Towards the Next Generation of Pre-Ranking System
- What Does It Take to Build a Performant Selective Classifier?
- AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- XGen-Q: An Explainable Domain-Adaptive LLM Framework with Retrieval-Augmented Generation for Software Security
- Learning to play: A Multimodal Agent for 3D Game-Play
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- Posits and the state of numerical representations in the age of exascale and edge computing
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- Reasoning-Enhanced Large Language Models for Molecular Property Prediction
- Weed Out, Then Harvest: Dual Low-Rank Adaptation is an Effective Noisy Label Detector for Noise-Robust Learning
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
- A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systems
- MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
- Triple-cooperative Video Shadow Detection
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
- Learning Representations Through Contrastive Neural Model Checking
- Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
- Randomized Matrix Sketching for Neural Network Training and Gradient Monitoring
- Generative AI for subgrid turbulence in large-eddy simulations
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
- Building the EHR Foundation Model via Next Event Prediction
- NTIRE 2020 Challenge on Spectral Reconstruction from an RGB Image
- Causally-Enhanced Reinforcement Policy Optimization
- Survey of Machine Learning Accelerators
- InfiR2: A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models
- Learning Admissible Heuristics for A*: Theory and Practice
- Compute-Optimal Quantization-Aware Training
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- Region-of-Interest Augmentation for Mammography Classification under Patient-Level Cross-Validation
- FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Accelerating Gravitational N-Body Simulations Using the RISC-V-Based Tenstorrent Wormhole
- k-Same-Siamese-GAN: k-Same Algorithm with Generative Adversarial Network for Facial Image De-identification with Hyperparameter Tuning and Mixed Precision Training
- TF-Replicator: Distributed Machine Learning for Researchers
- Elucidating the Design Space of FP4 training
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- GPU Temperature Simulation-Based Testing for In-Vehicle Deep Learning Frameworks
- Rethinking "Batch" in BatchNorm
- Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- Diffusion Models Beat GANs on Image Synthesis
- UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
- Chameleon: Taming Dynamic Operator Sequences for Memory-Intensive LLM Training
- Quantum parameter estimation with uncertainty quantification from continuous measurement data using neural network ensembles
- Training Multilingual Pre-trained Language Model with Byte-level Subwords
- Tri-Accel: Curvature-Aware Precision-Adaptive and Memory-Elastic Optimization for Efficient GPU Usage
- WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- Breaking the Conventional Forward-Backward Tie in Neural Networks: Activation Functions
- Compounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network
- MedLiteNet: Lightweight Hybrid Medical Image Segmentation Model
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- From Discord to Harmony: Decomposed Consonance-based Training for Improved Audio Chord Estimation
- Continuous Determination of Respiratory Rate in Hospitalized Patients using Machine Learning Applied to Electrocardiogram Telemetry
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer
- Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
- Backprop with Approximate Activations for Memory-efficient Network Training
- Automated classification of natural habitats using ground-level imagery
- High-Accuracy Low-Precision Training
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Formal Algorithms for Model Efficiency
- STAS: Spatio-Temporal Adaptive Computation Time for Spiking Transformers
- A Study of BFLOAT16 for Deep Learning Training
- High-Throughput Low-Cost Segmentation of Brightfield Microscopy Live Cell Images
- E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs
- Training DNNs with Hybrid Block Floating Point
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
- Latent Space Diffusion for Topology Optimization
- TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
- Deeper Inside Deep ViT
- FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose Estimation
- NTT's Machine Translation Systems for WMT19 Robustness Task
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
- Imbalance-Robust and Sampling-Efficient Continuous Conditional GANs via Adaptive Vicinal Learning and Auxiliary Regularization
- OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
- Compression-Induced Communication-Efficient Large Model Training and Inferencing
- Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
- SGEMM-cube: Emulating FP32 GEMM on Ascend NPUs Using FP16 Cube Units with Precision Recovery
- A Fast and Robust BERT-based Dialogue State Tracker for Schema-Guided Dialogue Dataset
- Capturing star formation activity from compressed photometric images of galaxies
Discussions
Related