The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource
2022/11/09 by Elliot Sollis, Abayomi Mosaku, Ala Abid +27 · 61 citations
Biochemistry, Genetics and Molecular Biology · #Epigenetics and DNA Methylation #Genetic Associations and Epidemiology #Genetic Syndromes and Imprinting
paper · pdf · doi:10.1093/nar/gkac1010
openalex publication_date 2022/11/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
The NHGRI-EBI GWAS Catalog (www.ebi.ac.uk/gwas) is a FAIR knowledgebase providing detailed, structured, standardised and interoperable genome-wide association study (GWAS) data to >200 000 users per year from academic research, healthcare and industry. The Catalog contains variant-trait associations and supporting metadata for >45 000 published GWAS across >5000 human traits, and >40 000 full P-value summary statistics datasets. Content is curated from publications or acquired via author submission of prepublication summary statistics through a new submission portal and validation tool. GWAS data volume has vastly increased in recent years. We have updated our software to meet this scaling challenge and to enable rapid release of submitted summary statistics. The scope of the repository has expanded to include additional data types of high interest to the community, including sequencing-based GWAS, gene-based analyses and copy number variation analyses. Community outreach has increased the number of shared datasets from under-represented traits, e.g. cancer, and we continue to contribute to awareness of the lack of population diversity in GWAS. Interoperability of the Catalog has been enhanced through links to other resources including the Polygenic Score Catalog and the International Mouse Phenotyping Consortium, refinements to GWAS trait annotation, and the development of a standard format for GWAS data.
Citations
Cited by
- Common clonal hematopoiesis driver mutations have disparate effects on macrophage cytokines, clonal expansion, and atherogenesis.
- Epistatic contributions to human traits via transcription factor mechanisms
- Whole-genome sequencing of 490,640 UK Biobank participants
- Trans-ancestry genome-wide study of depression identifies 697 associations implicating cell types and pharmacotherapies
- Genome-wide association studies of infant and toddler temperament in European and multi-ancestry populations
- Selective remodelling of the adipose niche in obesity and weight loss
- Genome-wide association meta-analysis of age at onset of walking in over 70,000 infants of European ancestry
- Consensus meta-analysis of genome-wide association studies for Alzheimer’s disease and related dementias
- “Be sustainable”: EOSC‐Life recommendations for implementation of FAIR principles in life science data handling
- Genome-wide analysis of brain age identifies 59 associated loci and unveils relationships with mental and physical health
- A Joint Promoterome-Proteome Atlas Highlights the Molecular Diversity of Human Skeletal Muscles
- Allele-specific genomics decodes gene targets and mechanisms of the non-coding genome
- Clinical genetic variation across Hispanic populations in the Mexican Biobank
- A global genetic interaction map of a human cell reveals conserved principles of genetic networks
- Precisely defining disease variant effects in CRISPR-edited single cells
- De novo protein-coding gene variants in developmental stuttering
- Comprehensive mapping of genetic variation at Epromoters reveals pleiotropic association with multiple disease traits
- Organism-Specific Sequence Motifs Link Ribosomal RNAs to Brain Disorders
- SEC16B identified as a regulator of VLDL secretion and lipid accumulation in the liver
- Hepatic SEC16B regulates lipid homeostasis by coordinating VLDL secretion and lipid droplet expansion
- Health outcomes after national acute sleep deprivation events among the American public
- Alcohol use disorder–associated gene FNDC4 alters glutamatergic and GABAergic neurogenesis in neural organoids
- Actively Protective Combinatorial Analysis: a Scalable Novel Method for Detecting Variants that Contribute to Reduced Disease Prevalence in High-Risk Individuals
- Genetic neurodevelopmental clustering and dyslexia
- The association of a polygenic lifespan score with the risk of common age-related diseases and mortality
- ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants
- Urobiota analysis and genome-wide association study in pediatric recurrent urinary tract infections and vesicoureteral reflux
- Variant effect predictors: a systematic review and practical guide
- A transcriptomic signature that predicts prehypertension in adolescence and higher systolic blood pressure in childhood
- Correcting for Genomic Inflation Leads to Loss of Power in Large‐Scale Genome‐Wide Association Study Meta‐Analysis
- Genome-wide associations spanning 194 in-hospital drug dosage change phenotypes highlight diverse genetic backgrounds in concurrent drug therapy
- Mapping disease loci to biological processes via joint pleiotropic and epigenomic partitioning
- GeneSetCart: assembling, augmenting, combining, visualizing, and analyzing gene sets
- A systematic analysis of the contribution of genetics to multimorbidity and comparisons with primary care data
- Kidney multiome-based genetic scorecard reveals convergent coding and regulatory variants
- The extended EA ModelSet—a FAIR dataset for researching and reasoning enterprise architecture modeling practices
- Tracing the Origins of Human Disease: A Phylogenomic Toolkit for Identifying Evolutionary Trade‐Offs
- Genetic variants associated with triglyceride metabolism and fasting triglyceride concentrations: a systematic review and a meta-analysis
- Multi-omics characterization of developing forebrain organoids unravels the dynamic molecular events of Rett syndrome pathogenesis
- MetaboAnalyst 6.0: towards a unified platform for metabolomics data processing, analysis and interpretation
- Phantom epistasis through the lens of genealogies
- Enabling Down Syndrome Research through a Knowledge Graph-Driven Analytical Framework
- Systematic Evaluation of Somatic Contamination in Germline Genomes
- Isolating the Genetic Component of Mania in Bipolar Disorder
- 2026 Heart Disease and Stroke Statistics: A Report of US and Global Data From the American Heart Association
- Beyond Single Variants: A Pathway-Based Approach to Explore the Genetic Basis of Memory
- Network toxicology and Mendelian randomization reveal pathogenic factors of monoethyl phthalate-induced thyroid cancer
- Lithium deficiency and the onset of Alzheimer’s disease
- BrainBridge Characterizes Key Factors affecting Alzheimer’s Disease and Associated Phenotypes
- Evaluating the effects of archaic protein-altering variants in living human adults
- Genome-wide association study of long COVID
- Long‐term cardiovascular risk in women with hypertensive disorders of pregnancy: Insights from polygenic risk scores
- Variant Set Distillation
- medicX-KG: A Knowledge Graph for Pharmacists' Drug Information Needs
- Diversity and scale: Genetic architecture of 2068 traits in the VA Million Veteran Program
- Rare de novo damaging DNA variants are enriched in attention-deficit/hyperactivity disorder and implicate risk genes
- A progeria syndrome links DNA hypermethylation to age-related pathology
- Sparse matrix factorization robust to sample sharing across GWAS reveals interpretable genetic components
- Genetic and multi‐omic risk assessment of Alzheimer's disease implicates core associated biological domains
- PhenotypeToGeneDownloaderR: automated multi-source retrieval and validation of phenotype-associated genes
- Unraveling the influence of childhood emotional support on adult aging: Insights from the UK Biobank
- Variant Classification Using Proteomics‐Informed Large Language Models Increases Power of Rare Variant Association Studies and Enhances Target Discovery
Related