2021/07/23 by Jessica Díaz, Díaz, Jessica, Jorge Eduardo Pérez Pérez +7
Computer Science · Decision Sciences · Mathematics · #Construction Project Management and Performance #D.2 #FOS: Computer and information sciences #Methodology (stat.ME) #Open Source Software Innovations #Software Engineering (cs.SE) #Software Engineering Research #Software Engineering Techniques and Practices #cs.SE #stat.ME
paper · pdf · doi:10.48550/arxiv.2107.11449
20 pages, 5 figures, 8 tables
arxiv created 2021/07/23 · openalex publication_date 2021/07/23 · arxiv updated 2021/07/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In recent years, the qualitative research on empirical software engineering that applies Grounded Theory is increasing. Grounded Theory (GT) is a technique for developing theory inductively e iteratively from qualitative data based on theoretical sampling, coding, constant comparison, memoing, and saturation, as main characteristics. Large or controversial GT studies may involve multiple researchers in collaborative coding, which requires a kind of rigor and consensus that an individual coder does not. Although many qualitative researchers reject quantitative measures in favor of other qualitative criteria, many others are committed to measuring consensus through Inter-Rater Reliability (IRR) and/or Inter-Rater Agreement (IRA) techniques to develop a shared understanding of the phenomenon being studied. However, there are no specific guidelines about how and when to apply IRR/IRA during the iterative process of GT, so researchers have been using ad hoc methods for years. This paper presents a process for systematically applying IRR/IRA in GT studies that meets the iterative nature of this qualitative research method, which is supported by a previous systematic literature review on applying IRR/RA in GT studies in software engineering. This process allows researchers to incrementally generate a theory while ensuring consensus on the constructs that support it and, thus, improving the rigor of qualitative research. This formalization helps researchers to apply IRR/IRA to GT studies when various raters are involved in coding. Measuring consensus among raters promotes communicability, transparency, reflexivity, replicability, and trustworthiness of the research.