vix.ing · top · new · best · stats · spec

Data Resource Profile: The Copenhagen Hospital Biobank (CHB)

2020/08/03 by Erik Sørensen, Lene Christiansen, Bartłomiej Wilkowski +18 · 3 citations
Biochemistry, Genetics and Molecular Biology · Medicine · #Biomedical Text Mining and Ontologies #Colorectal Cancer Screening and Detection #Ethics in Clinical Research

paper · pdf · doi:10.1093/ije/dyaa157

openalex publication_date 2020/08/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31

Abstract

Every year thousands of blood samples are drawn from patients admitted to Danish hospitals for diagnostic and treatment purposes. Traditionally, the leftover blood will be discarded after the required laboratory analyses are completed. However, such residual blood samples can be highly valuable for biomedical research, given the vast amount of clinical information these samples are often accompanied by—directly from the electronic patient record or through linkage to medical registers. To this end the Copenhagen Hospital Biobank (CHB) was established, with the overall aim to facilitate research in health and disease by enabling researchers’ access to a large resource of well-defined patient samples. The setup and running of CHB takes advantage of existing infrastructure at the Blood Banking facilities in the Capital Region of Denmark. The biobank thus includes leftover EDTA whole blood from samples drawn for blood type testing or red cell antibody screening. These analyses are estimated to be carried out for approximately 25% of all hospitalized patients and outpatients at Danish hospitals. The stored blood samples are mainly suitable for DNA extraction, and the biobank is accordingly considered a DNA biobank, even though whole blood is the initially stored material. The CHB collection and storage was initiated in February 2009 at the Blood Bank unit of Department of Clinical Immunology at Copenhagen University Hospital (Rigshospitalet), Denmark, and in early 2012 was expanded to also include leftover blood samples from the remaining general hospitals in the capital area of the Region. The CHB collection is ongoing, thus continuing to include new patient samples for as long as funding is available. There are no exclusion criteria except that each patient is included only once, i.e. repetitive blood sampling is currently not an option. At present (February 2020) the biobank counts ∼423 000 unique patients (Table 1, Figure 1). Consecutive inclusion of samples in the Copenhagen Hospital Biobank Consecutive inclusion of samples in the Copenhagen Hospital Biobank Number of samples collected and genotyped in Copenhagen Hospital Biobank as of February 2020 Number of samples collected and genotyped in Copenhagen Hospital Biobank as of February 2020 As a hospital-driven biobank with a relatively broad collection scheme, the CHB covers a wide range of disease areas, making it a unique resource for targeted selection of specific patient groups. Denmark has a long tradition for recording a variety of information on individual citizens in administrative registers, and linkage through the Danish Civil Registration System (CRS) can provide information from National Health registries which keep information on disease, health, diagnosis and treatment of all Danish residents.1,2 The Copenhagen Hospital Biobank thus offers an excellent platform for large-scale genetic studies with clinical research purposes. Following the red cell antibody screening or blood type routine testing of blood samples, the general procedure is to keep the leftover material at 4 °C for 1 week as back-up in case of the need for further testing. Provided any material remains at that time, the samples are then transferred for biobanking, if the individual is not already included in CHB. In some situations, this lag time can be longer before further processing takes place, due to logistic matters. Each EDTA whole-blood sample is mixed thoroughly and two aliquots of 850 µl are pipetted into 2 D barcode-labelled Thermo Fisher Matrix cryotubes, currently using the automatic Hamilton Star Liquid Handler (Hamilton Robotics). Previously, the same procedure was performed on a Biomek FX Liquid Handler (Beckman Coulther). One aliquot is stored at −20 °C in the automated sample storage system Brooks Universal Labstore (Brooks Life Science Systems). The other aliquot is stored at −80 °C in manual freezers for long-term storage, as reserve material. Due to the sampling scheme, where it may take a week or more before the samples are biobanked; the preferable downstream use of the samples is genetic analysis, as these pre-storage conditions have little impact on the integrity of DNA.3 DNA has been extracted from selected patient groups on a project-by-project basis. The majority of the DNA preparations were carried out by automated procedures using either Chemagic 360 (Perkin Elmer) or Hamilton Chemagic Star (Hamilton Robotics), and all genome-wide genetic analyses were exclusively performed on DNA extracted by the automated procedures. DNA preparations were routinely quantified (mean concentration = 56.6 ng/µl), and quality tested using standard spectrophotometric measures, followed by storage at −20°C until usage in genetic analysis. Since the source blood specimens follow a normal clinical flow, without any specific precaution measures taken before arriving at the biobank, there could potentially be at the least a minor risk for cross contamination between samples. To evaluate the extent of this risk, we performed a cross-contamination study on 560 EDTA whole-blood samples tested Rhesus D (RhD) negative by serological routine methods in the Blood Bank, and which had gone through the same procedures as the CHB-eligible blood samples. The experimental design relies on the fact that only ∼15% of all patients are RhD-negative, and thus any contamination transferred to an RhD-negative sample will most likely originate from a RhD-positive sample. DNA was extracted from the collected samples and tested for the presence of RHD exon 7 using a quantitative real-time polymerase chain reaction (PCR) assay. This analysis enables detection of admixture of RhD-positive DNA into RhD-negative DNA at a ratio of 1:1000.4 The study revealed no indication of cross-contamination of the 560 samples arising during the routine analyses in the Blood Bank. Considering the upper confidence limits of the data, it was estimated that a maximum of 1% of samples are contaminated above 1:1000. We therefore conclude that the CHB blood samples can be used for standard genetic methods. DNA extraction and subsequent genetic analyses have been carried out on samples from selected patients above 18 years of age, after approval of study protocols by the relevant authorities (see Ethical issues below). Of the blood samples deposited in CHB to date, genome-wide genotype data are available for ∼154 000 patients as of February 2020 (Table 1). Genotyping was performed at deCODE Genetics (Iceland) using the Infinium Global Screening Array (Illumina®). This array examines more than 660 000 common genetic markers, selected to maximize the imputation accuracy. The array is additionally particularly well suited for clinical research as it features more than 28 000 known disease associated markers [www.illumina.com]. Studies are under way, including ongoing genotyping of an additional ∼100 000 patient samples from CHB (Table 1). Genotype data are cleaned using standard quality control parameters, including filtering markers on genotype call rate, Hardy-Weinberg equilibrium and minor allele frequency, and filtering samples on genotype call rate, relatedness, sex mismatches and outlying heterozygosity. Imputation is then performed with long-range phasing and haplotype imputation using a population of Scandinavians as reference panel. After post-imputation quality control genotyped and imputed markers with a minor allele frequency, ≥0.01 are available for genetic analyses. At the outset there are no other data attached to the biobanked samples, except for the sample ID that links to the patients’ Civil Registration System number, which is a 10-digit personal identification number unique for all Danish residents. However, with the purpose of linking the CHB samples to medical data, particularly disease information, the CHB inventory is deposited under the umbrella of the Danish Biobank Register [www.danishnationalbiobank.com], a national collaboration keeping data from large biobanks from all over the country. The register can be accessed freely online and provides immediate knowledge of the number and type of available biological samples with specific International Classification of Diseases, revision 8 [ICD-8] or International Statistical Classification of Diseases and Related Health Problems, revision 10 [ICD-10] disease codes, albeit on an aggregated level [www.danishnationalbiobank.com/danish-biobank-register]. The disease information is obtained from linkage to the Danish National Patient Registry (NPR), which covers data from all hospitalizations and data on primary, secondary and supplementary diagnoses for all inpatients, outpatients and patients seen in emergency departments since 1977 (see Table 3 for examples of available data).5,Figure 2 gives a non-exhaustive overview of the current prevalence of several common and rarer diseases among CHB patients by February 2020. An extended list of disease prevalence among patients in CHB according to the main ICD-10 disease categories is shown in Table 2. Details about the data quality of NPR records, e.g. ‘positive predictive value’ (proportion of patients with a registered disease who truly have the disease) can be found in.5 Alternatively, sample IDs can also be linked directly to the electronic patient record thereby providing additional phenotypic detail, for example by text mining of the clinical notes.6 Prevalence of selected diseases among patients in the Copenhagen Hospital Biobank Prevalence of selected diseases among patients in the Copenhagen Hospital Biobank Disease prevalence among patients in Copenhagen Hospital Biobank according to ICD-10 disease categories ICD-10: International Statistical Classification of Diseases and Related Health Problems, revision 10. The 33 males with diagnoses within the ICD-10 category O00–O99 conceivably represent coding errors, or persons who underwent a legal sex change which includes change of Civil Registration System number. Disease prevalence among patients in Copenhagen Hospital Biobank according to ICD-10 disease categories ICD-10: International Statistical Classification of Diseases and Related Health Problems, revision 10. The 33 males with diagnoses within the ICD-10 category O00–O99 conceivably represent coding errors, or persons who underwent a legal sex change which includes change of Civil Registration System number. Selected Danish health registries and databases Data on primary, secondary and supplementary diagnoses as well as procedure-related complications Treatment data include information on, for example, surgery, dialysis, cancer treatment, stem cell transplantations, intensive care, in-hospital medical treatment (e.g. chemotherapy) Examinations cover, for example, information on angiography, ultrasound scans, psychological evaluations etc Data on primary, secondary and supplementary diagnoses as well as procedure-related complications Treatment data include information on, for example, surgery, dialysis, cancer treatment, stem cell transplantations, intensive care, in-hospital medical treatment (e.g. chemotherapy) Examinations cover, for example, information on angiography, ultrasound scans, psychological evaluations etc Selected Danish health registries and databases Data on primary, secondary and supplementary diagnoses as well as procedure-related complications Treatment data include information on, for example, surgery, dialysis, cancer treatment, stem cell transplantations, intensive care, in-hospital medical treatment (e.g. chemotherapy) Examinations cover, for example, information on angiography, ultrasound scans, psychological evaluations etc Data on primary, secondary and supplementary diagnoses as well as procedure-related complications Treatment data include information on, for example, surgery, dialysis, cancer treatment, stem cell transplantations, intensive care, in-hospital medical treatment (e.g. chemotherapy) Examinations cover, for example, information on angiography, ultrasound scans, psychological evaluations etc Individual-level data from NPR can be made available for approved research projects (see Data access below) along with required demographic and clinical data from other relevant National Health Registries and Clinical Quality Databases. Data on various socioeconomic factors can be obtained from numerous administrative registers that are available through ‘Statistics Denmark’ [www.dst.dk/en], including for instance data on education, salaries and income, employment status, pensions, housing etc.2 In addition, data on laboratory test results may be available from the Danish Clinical Laboratory Information System,7 whereas typically used lifestyle variables like weight, body mass index, smoking habits and alcohol consumption, are often available for clinical samples (like the CHB samples) from the almost 100 different Clinical Quality Databases8,9 or the electronic patient record.6 A short list of key registries that are valuable for biomedical research is provided in Table 3.5,8–12 All genetic data are stored and handled in a protected, secure private cloud environment on a dedicated section of the 19 000 core Danish National Supercomputer for Life Sciences Computerome [www.computerome.dk]. CHB is classified as a ‘biobank for future research’ and the establishment was approved by the Danish Data Protection Agency (file number 2007-580015; institutional file number RH2007-30-4129/ I-suite 00678). Since the biological samples stored in CHB are leftover material from routine blood analyses, the patients are not asked for informed consent before inclusion. However, patients are informed about the opt-out possibility to have their biological specimens excluded from use in research in general. Thus, since 2004 a national Register on Tissue Application (Vævsanvendelsesregistret) lists all individuals who have chosen to opt out and whose samples cannot be used for research purposes. Before initiating any study within the frames of CHB, the Register on Tissue Application is consulted to ensure that no ineligible patients are included. Also, as a rule, genome-wide genetic analyses are not performed on samples drawn from patients under the age of 18 years at the time of blood sampling. In addition, all health research projects using biological material require approval from the Danish Health Research Ethics Committee System. During this procedure, dispensation for the collection of informed consent from the included patients may be granted by the Committee, if deemed appropriate. In a recent study, a total of 9500 DNA samples from CHB were initially selected for replication purposes in a large Icelandic case-control study of diverticular disease (5426 Icelandic cases) and diverticulitis (2764 Icelandic cases). Following a genome-wide sequencing effort in the Icelandic patients, 16 top hits were followed-up using a final replication cohort of 5970 Danish diverticular disease patients and 3020 Danish controls. When combining the two datasets, the study identified three loci which were associated with either diverticular disease or diverticulitis.13 In another case-control study, DNA samples from 586 carefully selected cervical cancer patients within CHB were used to evaluate the effect of two common filaggrin gene mutations on the risk of cervical cancer. Through linkage to health registries, the authors restricted inclusion to patients for whom first discharge diagnosis in the National Patient Registry was verified in the Danish Cancer Register. The study suggested that filaggrin gene mutation carriers among cases had a worse prognosis than non-carrier cases, although the mutations did not increase risk of cervical cancer per se.14 Currently, several large studies focusing on the genetics of cardiovascular disease, degenerative diseases including arthritis, and reproductive health, have been initiated. Each of these is conducted as a targeted selection of patients, which are identified through linkage to the health registries, followed by more detailed phenotyping and genome-wide genotyping. The central strength of the CHB is the possibility to pick out patients with broad or very particular diagnoses of common or rare diseases, from a very large and ever-increasing collection of patient blood samples. In addition, the collection of clinical routine blood samples is a cost-effective biobanking practice, thereby facilitating the continued increase. In combination with the demographic and medical information that can be retrieved by linkage to the comprehensive Danish registry system, this large biobank collection offers eminent possibilities for biomedical research. Thus, an additional strength of the biobank is that once the patients have been included in the biobank, they can be followed over time in national registries and other clinical databases and their data used in current or future studies of genetic predisposition to various diseases or response to treatment. A possible limitation of a biobank collection of leftover material is that only patients that are submitted to blood typing or red cell antibody screening are included. The majority of these patients may well be admitted for an intentional surgical treatment or oncology therapy, and therefore will indeed constitute a selected group of predominantly severely affected patients at risk of needing blood transfusion. The opportunity for classification of disease categories using the health registries will, however, diminish the adverse significance of this potential issue. Another limitation is that patients are only included in the biobank on one occasion, thus studies of change over time, like for instance epigenetic changes, are not feasible. Data from CHB can be accessed for use in research by both public and private sectors through collaboration. The CHB is open for collaboration with all interested parties under the following terms: an application must be prepared in collaboration with CHB and relevant clinicians from the Capital Region. An application form can be attained by contacting: [[email protected]]. The project subsequently needs clearance from the Copenhagen Hospital Biobank Committee to avoid overlap with existing research efforts. In addition, according to the Danish legislation, all projects should be notified to and receive permission from the Danish Health Research Ethics Committee System, as well as from Knowledge Centre on Data Protection Compliance, the current personal data protection supervision authority in the Capital Region of Denmark. In cases where direct linking to electronic patient records is desirable, access must also be approved by the Danish Patient Safety Authority. The Copenhagen Hospital Biobank was supported by the Novo Nordisk Foundation (grant numbers NNF14CC0001 and NNF17OC0027594) and by Department of Clinical Immunology, Copenhagen University Hospital. The establishment and development of the Danish Biobank Register system is supported by the Novo Nordisk Foundation (grant numbers 2010-11-12 and 2009-07-28 to Statens Serum Institut to establish the Danish National Biobank). We are grateful to staff at all departments of clinical immunology in the general hospitals of the capital area, as well as the staff in the Biobank unit of Copenhagen University Hospital, for their assistance in establishment and operation of the Copenhagen Hospital Biobank. Authors affiliated with deCODE genetics declare competing interests as employees. All other authors declare no competing interests.

Cited by

Related