vix.ing · top · new · best · stats · spec

On the Time-Based Conclusion Stability of Cross-Project Defect\n Prediction Models

2019/11/14 by Abdul Ali Bangash, Hareem Sahar, Bangash, Abdul Ali +5
Computer Science · #Software Engineering Research #Software Reliability and Analysis Research #Software System Performance and Reliability

paper · pdf · doi:10.48550/arxiv.1911.06348

Abstract

Researchers in empirical software engineering often make claims based on\nobservable data such as defect reports. Unfortunately, in many cases, these\nclaims are generalized beyond the data sets that have been evaluated. Will the\nresearcher's conclusions hold a year from now for the same software projects?\nPerhaps not. Recent studies show that in the area of Software Analytics,\nconclusions over different data sets are usually inconsistent. In this article,\nwe empirically investigate whether conclusions in the area of defect prediction\ntruly exhibit stability throughout time or not. Our investigation applies a\ntime-aware evaluation approach where models are trained only on the past, and\nevaluations are executed only on the future. Through this time-aware\nevaluation, we show that depending on which time period we evaluate defect\npredictors, their performance, in terms of F-Score, the area under the curve\n(AUC), and Mathews Correlation Coefficient (MCC), varies and their results are\nnot consistent. The next release of a product, which is significantly different\nfrom its prior release, may drastically change defect prediction performance.\nTherefore, without knowing about the conclusion stability, empirical software\nengineering researchers should limit their claims of performance within the\ncontexts of evaluation, because broad claims about defect prediction\nperformance might be contradicted by the next upcoming release of a product\nunder analysis.\n

Related