From Tea Leaves to System Maps: A Survey and Framework on Context-aware Machine Learning Monitoring
2025/06/12 by Leest, Joran, Raibulet, Claudia, Lago, Patricia +1 · 2 citations
#FOS: Computer and information sciences #Software Engineering (cs.SE)
paper · doi:10.48550/arxiv.2506.10770
Abstract
Machine learning (ML) models in production fail when their broader systems -- from data pipelines to deployment environments -- deviate from training assumptions, not merely due to statistical anomalies in input data. Despite extensive work on data drift, data validation, and out-of-distribution detection, ML monitoring research remains largely model-centric while neglecting contextual information: auxiliary signals about the system around the model (external factors, data pipelines, downstream applications). Incorporating this context turns statistical anomalies into actionable alerts and structured root-cause analysis. Drawing on a systematic review of 94 primary studies, we identify three dimensions of contextual information for ML monitoring: the system element concerned (natural environment or technical infrastructure); the aspect of that element (runtime states, structural relationships, prescriptive properties); and the representation used (formal constructs or informal formats). This forms the Contextual System-Aspect-Representation (C-SAR) framework, a descriptive model synthesizing our findings. We identify 20 recurring triplets across these dimensions and map them to the monitoring activities they support. This study provides a holistic perspective on ML monitoring: from interpreting "tea leaves" (i.e., isolated data and performance statistics) to constructing and managing "system maps" (i.e., end-to-end views that connect data, models, and operating context).
Citations
- "We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine Learning
- On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- ML-On-Rails: Safeguarding Machine Learning Models in Software Systems A Case Study
- Designing monitoring strategies for deployed machine learning algorithms: navigating performativity through a causal lens
- Test & Evaluation Best Practices for Machine Learning-Enabled Systems
- Out of Distribution Detection via Domain-Informed Gaussian Process State Space Models
- Monitoring Algorithmic Fairness under Partial Observations
- FeedbackLogs: Recording and Incorporating Stakeholder Feedback into Machine Learning Pipelines
- Causal fault localisation in dataflow systems
- Deployment of Image Analysis Algorithms under Prevalence Shifts
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution Shifts
- Two-stage Modeling for Prediction with Confidence
- Estimating and Explaining Model Performance When Both Covariates and Labels Shift
- Temporal quality degradation in AI models
- Agreement-on-the-Line: Predicting the Performance of Neural Networks under Distribution Shift
- A Human-Centric Take on Model Monitoring
- Perspectives on Incorporating Expert Feedback into Model Updates
- Monitoring AI systems: A Problem Analysis, Framework and Outlook
- Machine Learning Operations (MLOps): Overview, Definition, and Architecture
- Data Smells: Categories, Causes and Consequences, and Detection of Suspicious Data in AI-based Systems
- Leveraging Unlabeled Data to Predict Out-of-Distribution Performance
- Contrastive Identification of Covariate Shift in Image Data
- Predicting with Confidence on Unseen Distributions
- Mandoline: Model Evaluation under Distribution Shift
- MLDemon: Deployment Monitoring for Machine Learning Systems
- Human-in-the-loop Handling of Knowledge Drift
- On the experiences of adopting automated data validation in an industrial machine learning project
- Why did the distribution change?
- Challenges in Deploying Machine Learning: A Survey of Case Studies
- Overton: A Data System for Monitoring and Improving Machine-Learned\n Products
- Machine Learning Testing: Survey, Landscapes and Horizons
- Machine Learning Testing: Survey, Landscapes and Horizons
- Identifying, categorizing and mitigating threats to validity in software engineering secondary studies
- Runaway Feedback Loops in Predictive Policing
Cited by
Related