Inference and missing data
1976/01/01 by Donald B. Rubin, DONALD B. RUBIN · 204 citations
Mathematics · #Statistical Methods and Bayesian Inference #Advanced Statistical Methods and Models #Statistical Methods and Inference
paper · doi:10.1093/biomet/63.3.581
Abstract
When making sampling distribution inferences about the parameter of the data, θ, it is appropriate to ignore the process that causes missing data if the missing data are ‘missing at random’ and the observed data are ‘observed at random’, but these inferences are generally conditional on the observed pattern of missing data. When making direct-likelihood or Bayesian inferences about θ, it is appropriate to ignore the process that causes missing data if the missing data are missing at random and the parameter of the missing data process is ‘distinct’ from θ. These conditions are the weakest general conditions under which ignoring the process that causes missing data always leads to correct inferences.
Cited by
- Establishment of VCA and EBNA1 IgA‐based combination by enzyme‐linked immunosorbent assay as preferred screening method for nasopharyngeal carcinoma: a two‐stage design with a preliminary performance study and a mass screening in southern China
- To fill or not to fill: Comparing imputation methods for improved riverine long‐term biodiversity monitoring
- Efficient Estimation of Average Treatment Effects Using the Estimated Propensity Score
- Fitting Ordinal Factor Analysis Models With Missing Data: A Comparison Between Pairwise Deletion and Multiple Imputation
- The Assessment of Reliability Under Range Restriction
- Evaluating Close Fit in Ordinal Factor Analysis Models With Multiply Imputed Data
- Distinguishing “Missing at Random” and “Missing Completely at Random”
- A Note on Bayesian Inference After Multiple Imputation
- Behaviour–Affect Pathways and Sociodemographic Moderation in the Well-being Mechanism of Urban Parks: Evidence From Yangzhou, a ‘Park City’
- AI solutions for evolutionary genomics of nonmodel species
- Marginal Models for Correlated Binary Responses with Multiple Classes and Multiple Levels of Nesting
- The Relationship Between Dropout and Outcome in Naturalistic Cognitive Behavior Therapy
- Missing Data Techniques for Structural Equation Modeling.
- Recent Developments in the Econometrics of Program Evaluation
- Family processes as pathways from income to young children's development.
- Connecting high school physics experiences, outcome expectations, physics identity, and physics career choice: A gender study
- The long-term financial outcome of children diagnosed with ADHD.
- Graphical Causal Models for Survey Inference
- Phylogenetically informed predictions outperform predictive equations in real and simulated data
- On Structural Equation Modeling with Data that are not Missing Completely at Random
- The fallacy of single imputation for trait databases: Use multiple imputation instead
- Logarithmic imputation techniques for temporal surveys: a memory-based approach explored through simulation and real-life applications
- Investigating Ceiling Effects in Longitudinal Data Analysis
- A Mixture Response Time Process Model for Aberrant Behaviors and Item Nonresponses
- Evaluating Supplemental Samples in Longitudinal Research: Replacement and Refreshment Approaches
- Applying the Bollen-Stine Bootstrap for Goodness-of-Fit Measures to Structural Equation Models with Missing Data
- Structure and Stress: Trajectories of Depressive Symptoms across Adolescence and Young Adulthood
- Associations among Parental Phubbing, Self-esteem, and Adolescents’ Proactive and Reactive Aggression: A Three-Year Longitudinal Study in China
- Explaining the temporal and spatial dimensions of robbery: Differences across measures of the physical and social environment
- The Treatment of Missing Data in Multivariate Analysis
- A Unified Approach to Measurement Error and Missing Data: Overview and Applications
- A Comparison of Three Popular Methods for Handling Missing Data: Complete-Case Analysis, Inverse Probability Weighting, and Multiple Imputation
- Nonparametric Tests of Panel Conditioning and Attrition Bias in Panel Surveys
- An Optimal Stratification Method for Addressing Nonresponse Bias in Bayesian Adaptive Survey Design
- Missing Data Analysis: Making It Work in the Real World
- Imputation methods of missing data for estimating the population mean using simple random sampling with known correlation coefficient
- A Practical Guide to Counterfactual Estimators for Causal Inference with Time‐Series Cross‐Sectional Data
- A Population-Based Study of Sexual Orientation Identity and Gender Differences in Adult Health
- A hierarchical latent response model for inferences about examinee engagement in terms of guessing and item‐level non‐response
- Inference, Learning, and Population Size: Projectivity for SRL Models
- Statistical Inference after Kernel Ridge Regression Imputation under item nonresponse
- Multiple Imputation for Longitudinal Data: A Tutorial
- Estimating spatially varying health effects of wildland fire smoke using mobile health data
- A Mixed-effects Model for Incomplete Data With Batch-Level Abundance-Dependent Missing-Data Mechanism
- Statistical paradises and paradoxes in big data (I): Law of large populations, big data paradox, and the 2016 US presidential election
- Missing data in single-cell transcriptomes reveals transcriptional shifts
- The Role of Sampling Weights When Modeling Survey Data
- Using simulation studies to evaluate statistical methods
- Deductive semiparametric estimation in Double-Sampling Designs with application to PEPFAR
- Effects of Multi-Aspect Online Reviews with Unobserved Confounders: Estimation and Implication
- Bayesian Inference from Non-Ignorable Network Sampling Designs
- Causal Inference: A Missing Data Perspective
- Handling Missingness, Failures, and Non-Convergence in Simulation Studies: A Review of Current Practices and Recommendations
- GRB Redshift Classifier to Follow-up High-Redshift GRBs Using Supervised Machine Learning
- Recurrent Neural Networks for Multivariate Time Series with Missing Values
- Robustness of Generalized Estimating Equation (GEE) Tests of Significance against Misspecification of the Error Structure Model
- Towards A Pan–Cultural Personality Structure: Input from 11 Psycholexical Studies
- Towards Generating Real-World Time Series Data
- Predicting Poverty
- Robust Optimal Designs when Missing Data Happen at Random
- Collaborative Filtering and the Missing at Random Assumption
- Semiparametric Regression for Repeated Outcomes with Nonignorable Nonresponse
- Biased-sample empirical likelihood weighting: an alternative to inverse probability weighting
- not-MIWAE: Deep Generative Modelling with Missing not at Random Data
- Imputation of Missing Data Using Linear Gaussian Cluster-Weighted\n Modeling
- Testing Moderation in Business and Psychological Studies with Latent Moderated Structural Equations
- Imputation by power transformation
- A Groupwise Approach for Inferring Heterogeneous Treatment Effects in Causal Inference
- Strengthening the Reporting of Observational Studies in Epidemiology (STROBE)
- Doubly Robust Inference With Nonprobability Survey Samples
- On missing label patterns in semi-supervised learning
- Sequentially additive nonignorable missing data modeling using auxiliary marginal information
- Modern Dimension Reduction
- Robust Doubly Protected Estimators for Quantiles with Missing Data
- Causal Inference Using Potential Outcomes
- Model-assisted inference for treatment effects using regularized calibrated estimation with high-dimensional data
- Equipping the Offline Population with Internet Access in an Online Panel: Does It Make a Difference?
- Toward a standardized evaluation of imputation methodology
- Multilevel Modelling of Complex Survey Data
- The Use of Propensity Scores to Assess the Generalizability of Results from Randomized Trials
- Multiple-Bias Modelling for Analysis of Observational Data
- Analyses using multiple imputation need to consider missing data in auxiliary variables
- Network Model-Assisted Inference from Respondent-Driven Sampling Data
- Does exposure to artificial light in the morning reduce reaction time variability during cognitive control processing?
- Initial effectiveness of an ICBT-protocol for GAD in psychiatric care – A feasibility-pilot study
- When to Impute? Imputation before and during cross-validation
- The Strain From Procedural Injustice on Parolees: Bridging Procedural Justice Theory and General Strain Theory
- Scalable Online Survey Framework: from Sampling to Analysis
- Computational Strategies for Multivariate Linear Mixed-Effects Models With Missing Values
- An introduction to modern missing data analyses
- Should data ever be thrown away? Pooling interval-censored data sets with different precision
- Handling Missing Data in Decision Trees: A Probabilistic Approach
- Nonparametric Statistical Inference and Imputation for Incomplete Categorical Data
- Volunteer Work and Hedonic, Eudemonic, and Social Well‐Being
- Spilt conformal prediction with missing response
- PARENTAL PERCEPTIONS OF LEARNING LOSS DURING COVID-19 SCHOOL CLOSURES IN 2020
- On the consistency of supervised learning with missing values
- Screening for stratification in two-phase ('two- stage') epidemiological surveys
- Red and Processed Meat and Mortality in a Low Meat Intake Population
- Sensitivity Analysis for Unmeasured Confounding in Coarse Structural Nested Mean Models
- A semiparametric inference to regression analysis with missing covariates in survey data
- A Brief Guide to Structural Equation Modeling
- Handling missing data in clinical research
- Powerful extreme phenotype sampling designs and score tests for genetic association studies
- Weighted generalized estimating equations and unified estimation for longitudinal data with nonmonotone missing data patterns
- A Review of Missing Data Reporting Practices in the Field of Internationalization of Higher Education
- Associations between psychedelic-related and meditation-related variables: A longitudinal study
- On the unnecessary ubiquity of hierarchical linear modeling.
- SAITS: Self-attention-based imputation for time series
- Clarifying Selection Bias in Cluster Randomized Trials: Estimands and Estimation
- An Evaluation of Dual Systems Theories of Adolescent Delinquency in a Normative Longitudinal Cohort Study of Youth
- Intra‐individual associations between intentional self‐regulation and prosocial behavior during adolescence: Evidence for bidirectionality
- Stereotype Promise: Racialized Teacher Appraisals of Asian American Academic Achievement
- A Bayesian Joint model for Longitudinal DAS28 Scores and Competing Risk Informative Drop Out in a Rheumatoid Arthritis Clinical Trial
- Relational and Individual Resources as Predictors of Empathy in Early Childhood
- Using causal diagrams to guide analysis in missing data problems
- Empirical Bayes Matrix Factorization
- A Cautious Note on Auxiliary Variables That Can Increase Bias in Missing Data Problems
- From contextual risk to preventive health behavior in a pandemic: a serial mediation analysis
- Using directed acyclic graphs to determine whether multiple imputation or subsample-multiple imputation estimates of an exposure-outcome association are unbiased
- Neuropathic Pain Diagnosis Simulator for Causal Discovery Algorithm Evaluation
- SMIM: a unified framework of Survival sensitivity analysis using Multiple Imputation and Martingale
- Three-step estimation of latent Markov models with covariates
- Recommendations for the Primary Analysis of Continuous Endpoints in Longitudinal Clinical Trials
- Inferring a Population Composition From Survey Data With Nonignorable Nonresponse: Borrowing Information From External Sources
- Estimating the causal effect of R&D subsidies in a pan-European program
- Shadow Rate Models and Monetary Policy
- Tough Shift: The Temporal Dynamics of Police Discretion
- Characterizing Mother‐Infant Dyadic Behaviors Following Infant Bids for Attention: Potential Mechanisms for Promoting Infant Attention Control and Language
- Socioeconomic and ethnic inequalities increase the risk of type 2 diabetes: an analysis of NHS health check attendees in Birmingham
- Causal effect estimates of online e-cigarette marketing exposure on future e-cigarette harm perception and use
- Assumptions and analysis planning in studies with missing data in multiple variables: moving beyond the MCAR/MAR/MNAR classification
- Group-based trajectory modeling under non-random attrition: A sensitivity analysis and application to frailty trajectories
- How do formal and informal science learning experiences during high school shape students’ career interest and STEM identity?
- Missing Data Sensitivity Analyses for Alcohol Research
- Handling missing data in modelling quality of clinician-prescribed routine care: Sensitivity analysis of departure from missing at random assumption
- Incomplete hierarchical data
- ESA CCI Soil Moisture GAPFILLED: an independent global gap-free satellite climate data record with uncertainty estimates
- Matrix Completion for Survey Data Prediction with Multivariate Missingness
- Self-harming behaviors among forensic psychiatric patients who committed violent offences: an exploratory study on the role of circumstances during the index offence and victim characteristics
- Online Political Participation of Refugees in Germany: Analysis of a Survey in Bavaria
- Model-Based Manifest and Latent Composite Scores in Structural Equation Models
- Modeling Missing Response Data in Item Response Theory: Addressing Missing Not at Random Mechanism with Monotone Missing Characteristics
- <b>mice</b>: Multivariate Imputation by Chained Equations in<i>R</i>
- How To Treat Missing Data In Survey Research
- Using Auxiliary Marginal Distributions in Imputations for Nonresponse while Accounting for Survey Weights, with Application to Estimating Voter Turnout
- Thinking Clearly About Sampling and Representation With the Total Survey Error Framework
- Maximum sampled conditional likelihood for informative subsampling
- What's a good imputation to predict with missing values?
- Semiparametric Optimal Estimation With Nonignorable Nonresponse Data
- Obtaining Causal Information by Merging Datasets with MAXENT
- The correlation-assisted missing data estimator
- Physician–patient racial concordance and disparities in birthing mortality for newborns
- A practical guide to causal discovery with cohort data
- Estimation of Regression Coefficients When Some Regressors are not Always Observed
- Prenatal exposure to per- and polyfluoroalkyl substances (PFAS) and their influence on inflammatory biomarkers in pregnancy: Findings from the LIFECODES cohort
- Itemwise conditionally independent nonresponse modeling for incomplete\n multivariate data
- An emotional regulation approach to psychosis recovery: The Living Through Psychosis group programme
- Adjusting for treatment effects in studies of quantitative traits: antihypertensive therapy and systolic blood pressure
- Popularity Bias in Recommendation: A Multi-stakeholder Perspective
- A Multisite Cluster Randomized Field Trial of Open Court Reading
- Sampling Techniques for Big Data Analysis
- What made you do this? Understanding black-box decisions with sufficient\n input subsets
- Nonresponse Bias Analysis in Longitudinal Studies: A Comparative Review with an Application to the Early Childhood Longitudinal Study
- Review: A gentle introduction to imputation of missing values
- Improving Missing Data Imputation with Deep Generative Models
- Analysis of Partially Observed Networks via Exponential-family Random Network Models
- Model Criticism for Bayesian Causal Inference
- Elected officials’ Online Sharing of Misinformation: Institutional and Ideological Checks
- Partial identification of mean achievement in ILSA studies with multi-stage stratified sample designs and student non-participation
- Home-based extended rehabilitation for older people with frailty (HERO): a multicentre randomised controlled trial with health economic analysis and process evaluation
- Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques
- An Apparent Paradox: A Classifier Trained from a Partially Classified Sample May Have Smaller Expected Error Rate Than That If the Sample Were Completely Classified
- Using Multiple Imputation to Classify Potential Outcomes Subgroups
- Likelihood inference for incompletely observed stochastic processes:\n ignorability conditions
- Generalizing trial findings using nested trial designs with sub-sampling of non-randomized individuals
- A Profile Likelihood Approach to Semiparametric Estimation with Nonignorable Nonresponse
- Learning representations for multivariate time series with missing data\n using Temporal Kernelized Autoencoders
- Large-scale Causal Approaches to Debiasing Post-click Conversion Rate Estimation with Multi-task Learning
- Estimating evolutionary parameters when viability selection is operating. [europepmc]
- Principled missing data methods for researchers. [europepmc]
- Tuning multiple imputation by predictive mean matching and local residual draws. [europepmc]
- Heart rate variability predicts levels of inflammatory markers: Evidence for the vagal anti-inflammatory pathway. [europepmc]
- Non-targeted UHPLC-MS metabolomic data processing methods: a comparative investigation of normalisation, missing value imputation, transformation and scaling. [europepmc]
- Testing measurement invariance in longitudinal data with ordered-categorical measures. [europepmc]
- EULAR/ACR classification criteria for adult and juvenile idiopathic inflammatory myopathies and their major subgroups: a methodology report. [europepmc]
- Characterizing and Managing Missing Structured Data in Electronic Health Records: Data Analysis. [europepmc]
- Efficacy of Dialectical Behavior Therapy for Adolescents at High Risk for Suicide: A Randomized Clinical Trial. [europepmc]
- Machine Learning and Integrative Analysis of Biomedical Big Data. [europepmc]
- Education and Cognitive Decline: An Integrative Analysis of Global Longitudinal Studies of Cognitive Aging. [europepmc]
- Ultra-processed food intake and risk of cardiovascular disease: prospective cohort study (NutriNet-Santé). [europepmc]
- Prognostic models for outcome prediction in patients with chronic obstructive pulmonary disease: systematic review and critical appraisal. [europepmc]
- Random forest-based imputation outperforms other methods for imputing LC-MS metabolomics data: a comparative study. [europepmc]
- Ultraprocessed Food Consumption and Risk of Type 2 Diabetes Among Participants of the NutriNet-Santé Prospective Cohort. [europepmc]
- Ethnic Differences in the Prevalence of Type 2 Diabetes Diagnoses in the UK: Cross-Sectional Analysis of the Health Improvement Network Primary Care Database. [europepmc]
- Effect of Doxycycline on Aneurysm Growth Among Patients With Small Infrarenal Abdominal Aortic Aneurysms: A Randomized Clinical Trial. [europepmc]
- Causality matters in medical imaging. [europepmc]
- Accuracy of random-forest-based imputation of missing data in the presence of non-normality, non-linearity, and interaction. [europepmc]
- A survey on missing data in machine learning. [europepmc]
- Association of Bariatric Surgery With Major Adverse Liver and Cardiovascular Outcomes in Patients With Biopsy-Proven Nonalcoholic Steatohepatitis. [europepmc]
- Comparison of Prolonged Exposure vs Cognitive Processing Therapy for Treatment of Posttraumatic Stress Disorder Among US Veterans: A Randomized Clinical Trial. [europepmc]
- Association of Bariatric Surgery With Cancer Risk and Mortality in Adults With Obesity. [europepmc]
- Effect of Verapamil on Pancreatic Beta Cell Function in Newly Diagnosed Pediatric Type 1 Diabetes: A Randomized Clinical Trial. [europepmc]
- Missing data in multi-omics integration: Recent advances through artificial intelligence. [europepmc]
Related