2014/07/08 by Russel Pears, Pears, Russel, Jacqui Finlay +3 · 1 citation
Computer Science · #Software Engineering Research #Software Reliability and Analysis Research #Imbalanced Data Classification Techniques
paper · pdf · doi:10.48550/arxiv.1407.2330
In this research we use a data stream approach to mining data and construct\nDecision Tree models that predict software build outcomes in terms of software\nmetrics that are derived from source code used in the software construction\nprocess. The rationale for using the data stream approach was to track the\nevolution of the prediction model over time as builds are incrementally\nconstructed from previous versions either to remedy errors or to enhance\nfunctionality. As the volume of data available for mining from the software\nrepository that we used was limited, we synthesized new data instances through\nthe application of the SMOTE oversampling algorithm. The results indicate that\na small number of the available metrics have significance for prediction\nsoftware build outcomes. It is observed that classification accuracy steadily\nimproves after approximately 900 instances of builds have been fed to the\nclassifier. At the end of the data streaming process classification accuracies\nof 80% were achieved, though some bias arises due to the distribution of data\nacross the two classes over time.\n