Onboard Optimization and Learning: A Survey
2025/05/07 by Pavel, Monirul Islam, Hu, Siyi, Pratama, Mahardhika +1
#FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2505.08793
Abstract
Onboard learning is a transformative approach in edge AI, enabling real-time data processing, decision-making, and adaptive model training directly on resource-constrained devices without relying on centralized servers. This paradigm is crucial for applications demanding low latency, enhanced privacy, and energy efficiency. However, onboard learning faces challenges such as limited computational resources, high inference costs, and security vulnerabilities. This survey explores a comprehensive range of methodologies that address these challenges, focusing on techniques that optimize model efficiency, accelerate inference, and support collaborative learning across distributed devices. Approaches for reducing model complexity, improving inference speed, and ensuring privacy-preserving computation are examined alongside emerging strategies that enhance scalability and adaptability in dynamic environments. By bridging advancements in hardware-software co-design, model compression, and decentralized learning, this survey provides insights into the current state of onboard learning to enable robust, efficient, and secure AI deployment at the edge.
Citations
- COMET: Co-Optimization of a CNN Model using Efficient-Hardware OBC Techniques
- Benchmarking Energy and Latency in TinyML: A Novel Method for Resource-Constrained AI
- FedHQ: Hybrid Runtime Quantization for Federated Learning
- Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services
- Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning
- Lightweight and Post-Training Structured Pruning for On-Device Large Lanaguage Models
- Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey
- Learning Critically: Selective Self Distillation in Federated Learning on Non-IID Data
- HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
- Membership Inference Attacks and Defenses in Federated Learning: A Survey
- Membership Inference Attacks and Defenses in Federated Learning: A Survey
- X-DFS: Explainable Artificial Intelligence Guided Design-for-Security Solution Space Exploration
- Vector Quantization Prompting for Continual Learning
- SpikeBottleNet: Spike-Driven Feature Compression Architecture for Edge-Cloud Co-Inference
- Fast On-device LLM Inference with NPUs
- Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
- Joint Pruning and Channel-wise Mixed-Precision Quantization for Efficient Deep Neural Networks
- Joint Pruning and Channel-Wise Mixed-Precision Quantization for Efficient Deep Neural Networks
- Low-Rank Quantization-Aware Training for LLMs
- Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference Systems
- Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
- Investigating White-Box Attacks for On-Device Models
- L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
- Splitwise: Efficient generative LLM inference using phase splitting
- LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning
- MirrorNet: A TEE-Friendly Framework for Secure On-device DNN Inference
- No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device ML
- Efficient Joint Optimization of Layer-Adaptive Weight Pruning in Deep Neural Networks
- A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- NormKD: Normalized Logits for Knowledge Distillation
- TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce Edge
- QLoRA: Efficient Finetuning of Quantized LLMs
- Efficient Coded Multi-Party Computation at Edge Networks
- Efficient Parallel Split Learning over Resource-constrained Wireless Edge Networks
- Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning
- Byzantine-Resilient Federated Learning at Edge
- FedLP: Layer-wise Pruning Mechanism for Communication-Computation Efficient Federated Learning
- X-Pruner: eXplainable Pruning for Vision Transformers
- Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators
- Workload-Balanced Pruning for Sparse Spiking Neural Networks
- A Comprehensive Survey of Continual Learning: Theory, Method and Application
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- Curriculum Temperature for Knowledge Distillation
- Designing and Training of Lightweight Neural Networks on Edge Devices using Early Halting in Knowledge Distillation
- SparCL: Sparse Continual Learning on the Edge
- On-Device Training Under 256KB Memory
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference
- A Survey on Computationally Efficient Neural Architecture Search
- Resource-Constrained Edge AI with Early Exit Prediction
- Knowledge Distillation from A Stronger Teacher
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
- Structured Pruning Learns Compact and Accurate Models
- A Fast Post-Training Pruning Framework for Transformers
- Overcoming Oscillations in Quantization-Aware Training
- Delta Keyword Transformer: Bringing Transformers to the Edge through Dynamically Pruned Multi-Head Self-Attention
- Model-Architecture Co-Design for High Performance Temporal GNN Inference on FPGA
- Membership Inference Attacks and Defenses in Neural Network Pruning
- Offloading Algorithms for Maximizing Inference Accuracy on Edge Device Under a Time Constraint
- DeepSteal: Advanced Model Extractions Leveraging Efficient Weight Stealing in Memories
- Decentralized Wireless Federated Learning with Differential Privacy
- AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric Learning
- DNN is not all you need: Parallelizing Non-Neural ML Algorithms on Ultra-Low-Power IoT Processors
- AutoFL: Enabling Heterogeneity-Aware Energy Efficient Federated Learning
- LoRA: Low-Rank Adaptation of Large Language Models
- MLPerf Tiny Benchmark
- Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
- Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
- Continual Learning via Bit-Level Information Preserving
- Post-training deep neural network pruning via layer-wise calibration
- The Power of Scale for Parameter-Efficient Prompt Tuning
- A Survey of Microarchitectural Side-channel Vulnerabilities, Attacks and Defenses in Cryptography
- Membership Inference Attacks on Machine Learning: A Survey
- HIR: An MLIR-based Intermediate Representation for Hardware Accelerator Description
- Learning Student-Friendly Teacher Networks for Knowledge Distillation
- Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
- DAQ: Channel-Wise Distribution-Aware Quantization for Deep Image Super-Resolution Networks
- CoEdge: Cooperative DNN Inference with Adaptive Workload Partitioning over Heterogeneous Edge Devices
- Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression
- ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural Networks
- Channel Pruning Guided by Spatial and Channel Attention for DNNs in Intelligent Edge Computing
- AdapterDrop: On the Efficiency of Adapters in Transformers
- Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture Search
- Gradient Flow in Sparse Neural Networks and How Lottery Tickets Win
- Online Continual Learning under Extreme Memory Constraints
- MCUNet: Tiny Deep Learning on IoT Devices
- Model Explanations with Differential Privacy
- Knowledge Distillation: A Survey
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
- Bayesian Bits: Unifying Quantization and Pruning
- SplitFed: When Federated Learning Meets Split Learning
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Training with Quantization Noise for Extreme Model Compression
- Dark Experience for General Continual Learning: a Strong, Simple Baseline
- Robust Pruning at Initialization
- Low-Complexity Distributed-Arithmetic-Based Pipelined Architecture for an LSTM Network
- iDLG: Improved Deep Leakage from Gradients
- Latent Replay for Real-Time Continual Learning
- Neural Network Pruning with Residual-Connections and Limited-Data
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Compacting, Picking and Growing for Unforgetting Continual Learning
- Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing
- Model Pruning Enables Efficient Federated Learning on Edge Devices
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- Deep Learning With Edge Computing: A Review
- Single Path One-Shot Neural Architecture Search with Uniform Sampling
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Class-incremental Learning via Deep Model Consolidation
- Improving Device-Edge Cooperative Inference of Deep Learning via 2-Step Pruning
- Federated Optimization in Heterogeneous Networks
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
- Cache Telepathy: Leveraging Shared Resource Attacks to Learn DNN Architectures
- DARTS: Differentiable Architecture Search
- Slalom: Fast, Verifiable and Private Execution of Neural Networks in Trusted Hardware
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- Adversarial Examples: Attacks and Defenses for Deep Learning
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
- Moonshine: Distilling with Cheap Convolutions
- Distributed Deep Neural Networks over the Cloud, the Edge and End Devices
- Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning
- BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
- Learning without Forgetting
Cited by
Related