vix.ing · top · new · best · stats

Matching IRT Models to Patient-Reported Outcomes Constructs: The Graded Response and Log-Logistic Models for Scaling Depression

2021/08/31 by Steven P. Reise, Han Du, Emily Wong +3 · 31 citations
Decision Sciences · Mathematics · Medicine · Psychology · #Child and Adolescent Psychosocial and Emotional Development #Clinical psychology #Cognition #Cognitive psychology #Computer science #Construct (python library) #Data mining #Econometrics #Item response theory #Logistic regression #Matching (statistics) #Maternal Mental Health During Pregnancy and Postpartum #Mathematics #Measure (data warehouse) #Multidimensional scaling #Psychiatry #Psychology #Psychometric Methodologies and Testing #Psychometrics #Scaling #Statistics

paper · pdf · doi:10.1007/s11336-021-09802-0

published in Psychometrika 86(3), 800-824 (Springer Science+Business Media)

openalex publication_date 2021/08/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04

Abstract

Item response theory (IRT) model applications extend well beyond cognitive ability testing, and various patient-reported outcomes (PRO) measures are among the more prominent examples. PRO (and like) constructs differ from cognitive ability constructs in many ways, and these differences have model fitting implications. With a few notable exceptions, however, most IRT applications to PRO constructs rely on traditional IRT models, such as the graded response model. We review some notable differences between cognitive and PRO constructs and how these differences can present challenges for traditional IRT model applications. We then apply two models (the traditional graded response model and an alternative log-logistic model) to depression measure data drawn from the Patient-Reported Outcomes Measurement Information System project. We do not claim that one model is "a better fit" or more "valid" than the other; rather, we show that the log-logistic model may be more consistent with the construct of depression as a unipolar phenomenon. Clearly, the graded response and log-logistic models can lead to different conclusions about the psychometrics of an instrument and the scaling of individual differences. We underscore, too, that, in general, explorations of which model may be more appropriate cannot be decided only by fit index comparisons; these decisions may require the integration of psychometrics with theory and research findings on the construct of interest.

Citations

Cited by

Related