MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
2015/12/03 by Tianqi Chen, Mu Li, Chen, Tianqi +17 · 275 citations
Computer Science · Engineering · #Advanced Memory and Neural Computing #Advanced Neural Network Applications #IoT and Edge/Fog Computing #cs.DC #cs.LG #cs.MS #cs.NE
paper · pdf · doi:10.48550/arxiv.1512.01274
In Neural Information Processing Systems, Workshop on Machine Learning Systems, 2016
arxiv created 2015/12/03 · arxiv updated 2015/12/07
Abstract
MXNet is a multi-language machine learning (ML) library to ease the development of ML algorithms, especially for deep neural networks. Embedded in the host language, it blends declarative symbolic expression with imperative tensor computation. It offers auto differentiation to derive gradients. MXNet is computation and memory efficient and runs on various heterogeneous systems, ranging from mobile devices to distributed GPU clusters. This paper describes both the API design and the system implementation of MXNet, and explains how embedding of both symbolic expression and tensor operation is handled in a unified fashion. Our preliminary experiments reveal promising results on large scale deep neural network applications using multiple GPU machines.
Cited by
- LFFD: A Light and Fast Face Detector for Edge Devices
- Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
- Bell Numbers and Stirling Numbers of the Mycielskian of Trees
- TD-Orch: Efficient Task-Data Orchestration for Distributed Systems with Application to Graph Processing
- Learning Where to Focus for Efficient Video Object Detection
- Online normalizer calculation for softmax
- A Proof of Useful Work for Artificial Intelligence on the Blockchain
- Benchmarking State-of-the-Art Deep Learning Software Tools
- Multiple Adaptive Bayesian Linear Regression for Scalable Bayesian Optimization with Warm Start
- Generative Low-bitwidth Data Free Quantization
- DGL-KE: Training Knowledge Graph Embeddings at Scale
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- EagerPy: Writing Code That Works Natively with PyTorch, TensorFlow, JAX, and NumPy
- Online Machine Learning in Big Data Streams
- Using Intuition from Empirical Properties to Simplify Adversarial Training Defense
- Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
- Fifer: Tackling Underutilization in the Serverless Era
- Attention as Activation
- Open Graph Benchmark: Datasets for Machine Learning on Graphs
- dxtb—An efficient and fully differentiable framework for extended tight-binding
- Revisiting Batch Normalization For Practical Domain Adaptation
- Training Deep Nets with Sublinear Memory Cost
- Bayesian Layers: A Module for Neural Network Uncertainty
- A Survey on Deep Learning Toolkits and Libraries for Intelligent User Interfaces
- Retention Time of Peptides in Liquid Chromatography Is Well Estimated upon Deep Transfer Learning
- Deep Frequent Spatial Temporal Learning for Face Anti-Spoofing
- Every Filter Extracts A Specific Texture In Convolutional Neural Networks
- FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads
- Demystifying Differentiable Programming: Shift/Reset the Penultimate Backpropagator
- A physics-informed deep learning framework for inversion and surrogate modeling in solid mechanics
- Array Programming with NumPy
- Spatial-Temporal Synchronous Graph Convolutional Networks: A New Framework for Spatial-Temporal Network Data Forecasting
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
- Moniqua: Modulo Quantized Communication in Decentralized SGD
- Adaptive Online Learning with Momentum for Contingency-based Voltage Stability Assessment
- Stanza: Layer Separation for Distributed Training in Deep Learning
- A Distributed Multi-GPU System for Large-Scale Node Embedding at Tencent
- Deep3D: Fully Automatic 2D-to-3D Video Conversion with Deep Convolutional Neural Networks
- CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning
- Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-grained Image Recognition
- SparCML: High-Performance Sparse Communication for Machine Learning
- Cavs: A Vertex-centric Programming Interface for Dynamic Neural Networks
- The Convergence of Sparsified Gradient Methods
- DeepLocalize: Fault Localization for Deep Neural Networks
- FanStore: Enabling Efficient and Scalable I/O for Distributed Deep Learning
- FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- Conditional Positional Encodings for Vision Transformers
- OneFlow: Redesign the Distributed Deep Learning Framework from Scratch
- DYNAMIX: RL-based Adaptive Batch Size Optimization in Distributed Machine Learning Systems
- Context-Aware Drive-thru Recommendation Service at Fast Food Restaurants
- Multi-Objective De Novo Drug Design with Conditional Graph Generative Model
- Tensor Regression Networks
- Diverse Sample Generation: Pushing the Limit of Generative Data-free Quantization
- Confidence Regularized Self-Training
- Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep Learning
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- MLPerf Training Benchmark
- From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures
- Demystifying Neural Style Transfer
- FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
- A deep learning framework for solution and discovery in solid mechanics
- A Survey on Serverless Computing
- A Network Structure to Explicitly Reduce Confusion Errors in Semantic Segmentation
- Seesaw-Net: Convolution Neural Network With Uneven Group Convolution
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems
- Hybrid Composition with IdleBlock: More Efficient Networks for Image Recognition
- DMLO: Deep Matching LiDAR Odometry
- Data-Driven Sparse Structure Selection for Deep Neural Networks
- ScaleFreeCTR: MixCache-based Distributed Training System for CTR Models with Huge Embedding Table
- A comprehensive study of batch construction strategies for recurrent neural networks in MXNet
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
- Hybrid Data-Model Parallel Training for Sequence-to-Sequence Recurrent Neural Network Machine Translation
- Multi-style Generative Network for Real-time Transfer
- Deep Learning At Scale and At Ease
- An Efficient Statistical-based Gradient Compression Technique for Distributed Training Systems
- TF-Replicator: Distributed Machine Learning for Researchers
- Fast, Better Training Trick -- Random Gradient
- Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs
- An Introduction to Deep Learning for the Physical Layer
- A Survey on Proactive Customer Care: Enabling Science and Steps to Realize it
- Interleaved Group Convolutions for Deep Neural Networks
- Structured Attentions for Visual Question Answering
- Real-Time Machine Learning: The Missing Pieces
- MediaPipe: A Framework for Building Perception Pipelines
- BackPACK: Packing more into backprop
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- Optimizer Fusion: Efficient Training with Better Locality and Parallelism
- The Neural Network Approach to Inverse Problems in Differential Equations
- SmartSwap: Swap-Based Memory Optimization for LLM Training under Varying Operator Sequences
- DyNet: The Dynamic Neural Network Toolkit
- DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Attentional Feature Fusion
- BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNet
- Invasiveness Prediction of Pulmonary Adenocarcinomas Using Deep Feature Fusion Networks
- Neural Network Distiller: A Python Package For DNN Compression Research
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Cellpose: a generalist algorithm for cellular segmentation
- Tensor Contraction Layers for Parsimonious Deep Nets
- A Runtime-Based Computational Performance Predictor for Deep Neural Network Training
- TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning
- Integral Human Pose Regression
- Deep Leakage from Gradients
- Deep Learning in Mobile and Wireless Networking: A Survey
- Boosting Model Performance through Differentially Private Model Aggregation
- 3D Context Enhanced Region-based Convolutional Neural Network for End-to-End Lesion Detection
- Ray: A Distributed Framework for Emerging AI Applications
- The AI gambit: leveraging artificial intelligence to combat climate change—opportunities, challenges, and recommendations
- ModelHub.AI: Dissemination Platform for Deep Learning Models
- DeLS-3D: Deep Localization and Segmentation with a 3D Semantic Map
- Priority-based Parameter Propagation for Distributed DNN Training
- Recent Advances and Applications of Deep Learning Methods in Materials Science
- ModiPick: SLA-aware Accuracy Optimization For Mobile Deep Inference
- MergeComp: A Compression Scheduler for Scalable Communication-Efficient Distributed Training
- Learning with a Strong Adversary
- Detection and Attention: Diagnosing Pulmonary Lung Cancer from CT by Imitating Physicians
- DISC: A Dynamic Shape Compiler for Machine Learning Workloads
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- RandomOut: Using a convolutional gradient norm to rescue convolutional filters
- Luandri: a Clean Lua Interface to the Indri Search Engine
- More Information Supervised Probabilistic Deep Face Embedding Learning
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
- A robust anomaly finder based on autoencoders
- Deep Factors with Gaussian Processes for Forecasting
- Sub-pixel face landmarks using heatmaps and a bag of tricks
- Graph Generative Models for Fast Detector Simulations in High Energy Physics
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Scheduling Optimization Techniques for Neural Network Training
- Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Face Detection with Feature Pyramids and Landmarks
- From Facial Expression Recognition to Interpersonal Relation Prediction
- A Survey on Deep Learning for Neuroimaging-based Brain Disorder Analysis
- OpenEI: An Open Framework for Edge Intelligence
- Sionnx: Automatic Unit Test Generator for ONNX Conformance
- Natural-Parameter Networks: A Class of Probabilistic Neural Networks
- Draw your Neural Networks
- DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks
- GPU-based Parallel Computation Support for Stan
- Vector representations of text data in deep learning
- dMath: A Scalable Linear Algebra and Math Library for Heterogeneous GP-GPU Architectures
- A First Look at Deep Learning Apps on Smartphones
- A temporal-to-spatial deep convolutional neural network for classification of hand movements from multichannel electromyography data
- Distributed Machine Learning through Heterogeneous Edge Systems
- A Survey on Large-scale Machine Learning
- DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
- Stochastic Distributed Optimization for Machine Learning from Decentralized Features
- Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
- Small, Accurate, and Fast Vehicle Re-ID on the Edge: the SAFR Approach
- On Improving Temporal Consistency for Online Face Liveness Detection
- Sinan: Data-Driven, QoS-Aware Cluster Management for Microservices
- Auto-STGCN: Autonomous Spatial-Temporal Graph Convolutional Network Search Based on Reinforcement Learning and Existing Research Results
- Deformable Tube Network for Action Detection in Videos
- Neural Networks for Lorenz Map Prediction: A Trip Through Time
- Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
- Dragon: A Computation Graph Virtual Machine Based Deep Learning Framework
- Domain Adaptation for Semantic Segmentation via Class-Balanced Self-Training
- Parallel Programming Models for Heterogeneous Many-Cores : A Survey
- Dynamic Key-Value Memory Networks for Knowledge Tracing
- Semantic Hierarchy Preserving Deep Hashing for Large-scale Image Retrieval
- A Good Practice Towards Top Performance of Face Recognition: Transferred Deep Feature Fusion
- Performance Evaluation of Deep Learning Tools in Docker Containers
- Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- FaceX-Zoo: A PyTorch Toolbox for Face Recognition
- TensorX: Extensible API for Neural Network Model Design and Deployment
- Intermittent Demand Forecasting with Deep Renewal Processes
- Characterizing the Deep Neural Networks Inference Performance of Mobile Applications
- Asynchronous Stochastic Proximal Methods for Nonconvex Nonsmooth Optimization
- Generative Poisoning Attack Method Against Neural Networks
- A Berkeley View of Systems Challenges for AI
- Fast and Accurate, Convolutional Neural Network Based Approach for Object Detection from UAV
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Distributed Training Large-Scale Deep Architectures
- Knowledge Projection for Deep Neural Networks
- An Efficient DP-SGD Mechanism for Large Scale NLP Models
- Revise Saturated Activation Functions
- Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
- SLSGD: Secure and Efficient Distributed On-device Machine Learning
- Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification
- Deep Hashing with Category Mask for Fast Video Retrieval
- Graph-Based Global Reasoning Networks
- Learning Deep Representations Using Convolutional Auto-encoders with Symmetric Skip Connections
- LAIF: AI, Deep Learning for Germany Suetterlin Letter Recognition and Generation
- RPC Considered Harmful: Fast Distributed Deep Learning on RDMA
- Bridging the Gap Between Neural Networks and Neuromorphic Hardware with A Neural Network Compiler
- A Framework for Democratizing AI
- Object Detection in Video with Spatial-temporal Context Aggregation
- DeepMarks: A Digital Fingerprinting Framework for Deep Neural Networks
- LPRNet: Lightweight Deep Network by Low-rank Pointwise Residual Convolution
- ByzShield: An Efficient and Robust System for Distributed Training
- Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks
- A Labeling-Free Approach to Supervising Deep Neural Networks for Retinal Blood Vessel Segmentation
- Demystifying the MLPerf Benchmark Suite
- Ripple: A Practical Declarative Programming Framework for Serverless Compute
- A survey on Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
- Scale-Aware Trident Networks for Object Detection
- Exascale Deep Learning for Scientific Inverse Problems
- Just ASK: Building an Architecture for Extensible Self-Service Spoken Language Understanding
- Hemingway: Modeling Distributed Optimization Algorithms
- HierTrain: Fast Hierarchical Edge AI Learning with Hybrid Parallelism in Mobile-Edge-Cloud Computing
- pCAMP: Performance Comparison of Machine Learning Packages on the Edges
- A Highly Configurable Hardware/Software Stack for DNN Inference Acceleration
- SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
- Frustrated with Replicating Claims of a Shared Model? A Solution
- Question Type Guided Attention in Visual Question Answering
- Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
- Face Detection with End-to-End Integration of a ConvNet and a 3D Model
- Parle: parallelizing stochastic gradient descent
- Revisit Batch Normalization: New Understanding from an Optimization View and a Refinement via Composition Optimization
- GraphTheta: A Distributed Graph Neural Network Learning System With Flexible Training Strategy
- Spectral Feature Transformation for Person Re-identification
- Omnivore: An Optimizer for Multi-device Deep Learning on CPUs and GPUs
- Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text Generation
- Deep Learning Training in Facebook Data Centers: Design of Scale-up and Scale-out Systems
- Progressive Neural Networks for Image Classification
- Data-driven forecasting of solar irradiance
- PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
- GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- ENT-DESC: Entity Description Generation by Exploring Knowledge Graph
- Towards Quantized Model Parallelism for Graph-Augmented MLPs Based on Gradient-Free ADMM Framework
- Graph-to-Sequence Learning using Gated Graph Neural Networks
- Deep Extreme Multi-label Learning
- Shuffle-Exchange Brings Faster: Reduce the Idle Time During Communication for Decentralized Neural Network Training
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- IGCV2: Interleaved Structured Sparse Convolutional Neural Networks
- Ludwig: a type-based declarative deep learning toolbox
- Stochastic Gradient MCMC with Stale Gradients
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- Symbolic Techniques for Deep Learning: Challenges and Opportunities
- Vanishing Nodes: Another Phenomenon That Makes Training Deep Neural Networks Difficult
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- SuperNeurons: FFT-based Gradient Sparsification in the Distributed Training of Deep Neural Networks
- Reinforcement Learning and Adaptive Sampling for Optimized DNN Compilation
- Stochastic Activation Pruning for Robust Adversarial Defense
- Probabilistic Forecasting with Temporal Convolutional Neural Network
- Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
- Lightweight Mask R-CNN for Long-Range Wireless Power Transfer Systems
- AI Enabling Technologies: A Survey
- DBLFace: Domain-Based Labels for NIR-VIS Heterogeneous Face Recognition
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- Hyperparameter Transfer Learning with Adaptive Complexity
- Scanner: Efficient Video Analysis at Scale
- Optimal Gradient Checkpoint Search for Arbitrary Computation Graphs
- MetaTune: Meta-Learning Based Cost Model for Fast and Efficient Auto-tuning Frameworks
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- Theano-MPI: a Theano-based Distributed Training Framework
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
- Advbox: a toolbox to generate adversarial examples that fool neural networks
- Network-accelerated Distributed Machine Learning Using MLFabric
- A Scalable and Cloud-Native Hyperparameter Tuning System
- Akid: A Library for Neural Network Research and Production from a Dataism Approach
- A Model Parallel Proximal Stochastic Gradient Algorithm for Partially Asynchronous Systems
- Parallel Training of Deep Networks with Local Updates
- GEVO: GPU Code Optimization using Evolutionary Computation
- Light-weighted Saliency Detection with Distinctively Lower Memory Cost and Model Size
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Slot Based Image Augmentation System for Object Detection
- STN-OCR: A single Neural Network for Text Detection and Text Recognition
- DASNet: Dynamic Activation Sparsity for Neural Network Efficiency Improvement
- AsymmNet: Towards ultralight convolution neural networks using asymmetrical bottlenecks
- HP-GNN: Generating High Throughput GNN Training Implementation on CPU-FPGA Heterogeneous Platform
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- MirBot: A collaborative object recognition system for smartphones using convolutional neural networks
- Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
Related