2025/01/01 by D J Dedman · 1 voice
Biochemistry, Genetics and Molecular Biology · Environmental Science · Neuroscience · #Genetic Associations and Epidemiology #Genetic Neurodegenerative Diseases #Health, Environment, Cognitive Aging
paper · pdf · doi:10.17037/pubs.04678099
openalex publication_date 2025/01/01 · openalex created_date 2026/01/21 · openalex updated_date 2026/07/01
BACKGROUND: Large databases of electronic health records (EHR) are an important resource for generating evidence on population patterns of health and disease, the safety and effectiveness of medicines, and the quality of health services. Combining data from multiple databases offers potential advantages over single database studies, including greater statistical power, and generalisability of results. A key challenge in multi-database studies is to extract comparable data from each source. Very few studies have combined data from two UK primary care databases to study exposure-disease associations and more evidence is needed on the strengths and limitations of different approaches for combining and analysing data in this specific context. OBJECTIVES: To review current practices and methods for combining data in multi-database studies. To compare and combine data from two UK primary care EHR databases – Clinical Practice Research Datalink (CPRD) GOLD and Aurum, using as a motivating example the reported association of cancer risk in patients with Huntington's disease (HD). METHORDS: I undertook a systematic review of multi-database studies using primary care EHR data, with a focus on how heterogeneity was investigated, reported and managed analytically. I compared trends in incidence and prevalence of cancer and HD in CPRD GOLD and Aurum primary care EHR databases, and undertook detailed comparisons of characteristics of patients with these conditions. Matched case-control and cohort studies were conducted to examine the association of cancer risk in patients with HD during premanifest and post diagnosis phases. RESULTS: The systematic review identified 109 multi-database studies, of which 62 examined one or more exposure-outcome associations. A one-stage pooled analysis was the most common 5 analytical approach. Assessment of and adjustments for heterogeneity or clustering were reported inconsistently. I defined groups of patients with HD or cancer in the two CPRD databases, and showed they were similar in respect of trends in incidence and prevalence, as well as other characteristics. Data curation methods may have contributed to differences in incidence and prevalence particularly during 1990-2000. In the case-control study newly diagnosed HD patients had a reduced risk of a prior cancer diagnosis compared to matched controls without HD (adjusted odds ratio 0.78, 95% confidence interval [CI] 0.59 - 1.04). In the cohort study cancer risk was lower among HD patients post diagnosis compared a matched group of patients without HD (adjusted hazard ratio 0.72, CI 0.54 – 0.96). CONCLUSIONS: This work contributes to the understanding of potential sources of heterogeneity in multidatabase studies with primary care EHR data. Differences in data curation methods and underlying data models present challenges to defining exactly equivalent clinical concepts in different databases. Researchers should be aware of these potential sources of variability when planning combined database studies and interpreting results. There were modest reductions in cancer risk among HD patients, but it was not possible to exclude negative surveillance bias as a possible explanation.