Scalability and Maintainability Challenges and Solutions in Machine Learning: Systematic Literature Review
2025/04/15 by Karthik Shivashankar, Shivashankar, Karthik, Ghadi S. Al Hajj +3 · 8 citations
Computer Science · #Machine Learning and Data Classification #Software Engineering Research #Software System Performance and Reliability
paper · pdf · doi:10.48550/arxiv.2504.11079
Abstract
This systematic literature review examines the critical challenges and solutions related to scalability and maintainability in Machine Learning (ML) systems. As ML applications become increasingly complex and widespread across industries, the need to balance system scalability with long-term maintainability has emerged as a significant concern. This review synthesizes current research and practices addressing these dual challenges across the entire ML life-cycle, from data engineering to model deployment in production. We analyzed 124 papers to identify and categorize 41 maintainability challenges and 13 scalability challenges, along with their corresponding solutions. Our findings reveal intricate inter dependencies between scalability and maintainability, where improvements in one often impact the other. The review is structured around six primary research questions, examining maintainability and scalability challenges in data engineering, model engineering, and ML system development. We explore how these challenges manifest differently across various stages of the ML life-cycle. This comprehensive overview offers valuable insights for both researchers and practitioners in the field of ML systems. It aims to guide future research directions, inform best practices, and contribute to the development of more robust, efficient, and sustainable ML applications across various domains.
Citations
- Machine/Deep Learning for Software Engineering: A Systematic Literature Review
- Kubric: A scalable dataset generator
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
- Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and Process
- Looper: An end-to-end ML platform for product decisions
- Towards a Common Testing Terminology for Software Engineering and Data Science Experts
- Pre-Trained Models: Past, Present and Future
- MLOps Challenges in Multi-Organization Setup: Experiences from Two Real-World Cases
- On the experiences of adopting automated data validation in an industrial machine learning project
- What Are We Really Testing in Mutation Testing for Machine Learning? A Critical Reflection
- Scalable federated machine learning with FEDn
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile Applications
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- A Software Engineering Perspective on Engineering Machine Learning Systems: State of the Art and Challenges
- A software engineering perspective on engineering machine learning systems: State of the art and challenges
- Communication optimization strategies for distributed deep neural network training: A survey
- Challenges in Deploying Machine Learning: a Survey of Case Studies
- Software engineering for artificial intelligence and machine learning software: A systematic literature review
- Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure
- MLCask: Efficient Management of Component Evolution in Collaborative Data Analytics Pipelines
- AI Lifecycle Models Need To Be Revised. An Exploratory Study in Fintech
- Machine Learning Systems in the IoT: Trustworthiness Trade-offs for Edge Intelligence
- High Performance Data Engineering Everywhere
- Adoption and Effects of Software Engineering Best Practices in Machine Learning
- Quality Management of Machine Learning Systems
- Kafka-ML: connecting the data stream with ML/AI frameworks
- Towards CRISP-ML(Q): A Machine Learning Process Model with Quality Assurance Methodology
- Callisto: Entropy based test generation and data quality assessment for Machine Learning Systems
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- Auptimizer -- an Extensible, Open-Source Framework for Hyperparameter Tuning
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- An Empirical Study towards Characterizing Deep Learning Development and Deployment across Different Frameworks and Platforms
- Overton: A Data System for Monitoring and Improving Machine-Learned Products
- Shuffler: A Large Scale Data Management Tool for ML in Computer Vision
- Machine Learning Testing: Survey, Landscapes and Horizons
- Machine Learning Testing: Survey, Landscapes and Horizons
- The Machine Learning Bazaar: Harnessing the ML Ecosystem for Effective System Development
- Data Cleaning for Accurate, Fair, and Robust Models: A Big Data - AI Integration Approach
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- On Testing Machine Learning Programs
- Software Engineering Challenges of Deep Learning
- Auto-Keras: An Efficient Neural Architecture Search System
- Tunability: Importance of Hyperparameters of Machine Learning Algorithms
- 500+ times faster than deep learning
- On the Pragmatic Design of Literature Studies in Software Engineering: An Experience-based Guideline
- Empirical studies of agile software development: A systematic review
- The Measurement of Observer Agreement for Categorical Data
Related