2021/08/11 by Mohammad Masudur Rahman, Rahman, Mohammad Masudur, Foutse Khomh +5 · 1 citation
Computer Science · #D.2 #D.2.5 #D.2.7 #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Software Engineering (cs.SE) #Software Engineering Research #Software System Performance and Reliability #Software Testing and Debugging Techniques
paper · pdf · doi:10.48550/arxiv.2108.05341
openalex publication_date 2021/08/11 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Being light-weight and cost-effective, IR-based approaches for bug\nlocalization have shown promise in finding software bugs. However, the accuracy\nof these approaches heavily depends on their used bug reports. A significant\nnumber of bug reports contain only plain natural language texts. According to\nexisting studies, IR-based approaches cannot perform well when they use these\nbug reports as search queries. On the other hand, there is a piece of recent\nevidence that suggests that even these natural language-only reports contain\nenough good keywords that could help localize the bugs successfully. On one\nhand, these findings suggest that natural language-only bug reports might be a\nsufficient source for good query keywords. On the other hand, they cast serious\ndoubt on the query selection practices in the IR-based bug localization. In\nthis article, we attempted to clear the sky on this aspect by conducting an\nin-depth empirical study that critically examines the state-of-the-art query\nselection practices in IR-based bug localization. In particular, we use a\ndataset of 2,320 bug reports, employ ten existing approaches from the\nliterature, exploit the Genetic Algorithm-based approach to construct optimal,\nnear-optimal search queries from these bug reports, and then answer three\nresearch questions. We confirmed that the state-of-the-art query construction\napproaches are indeed not sufficient for constructing appropriate queries (for\nbug localization) from certain natural language-only bug reports although they\ncontain such queries. We also demonstrate that optimal queries and non-optimal\nqueries chosen from bug report texts are significantly different in terms of\nseveral keyword characteristics, which has led us to actionable insights.\nFurthermore, we demonstrate 27%--34% improvement in the performance of\nnon-optimal queries through the application of our actionable insights to them.\n