vix.ing · top · new · best · stats · spec

A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays

2026/07/16 by Chang Min Yun, Salil Deshpande, Vivekanandan Ramalingam +7 · 1 voice
Biochemistry, Genetics and Molecular Biology · #Genomics and Chromatin Dynamics #Machine Learning in Bioinformatics #Chromatin Remodeling and Cancer #Chromatin #Sequence motif #ENCODE #Transcription factor #Motif (music) #Lexicon #DNA binding site #DNA #Binding site

paper · doi:10.5281/zenodo.21398817

openalex publication_date 2026/07/16 · openalex created_date 2026/07/18 · openalex updated_date 2026/07/19

Abstract

Abstract Regulatory DNA contains sequence elements that guide transcription factor (TF) binding and chromatin accessibility, that in turn control the expression of nearby genes. Most current descriptions of such regulatory elements are based on classical, statistical enrichment-based motif discovery methods applied to TF binding signals. Here, we present one of the largest collections of regulatory DNA deep learning models trained on TF binding and chromatin accessibility to date, and the first comprehensive lexicon of predictive regulatory motifs derived from the models. We trained BPNet and ChromBPNet models across 2,339 TF ChIP-seq datasets in 788 TFs, 1,143 DNase-seq datasets in 287 samples, and 369 ATAC-seq datasets in 204 samples. From the models, we extracted 286,836 sequence motifs that quantitatively predict regulatory signal in the cellular context of each dataset. To consolidate the discovered motifs into a single, non-redundant union across all cell contexts, we developed MotifCompendium—a GPU-accelerated motif management package, that can perform an accelerated calculation of motif-optimized pairwise similarity between 10,000 motifs in just 6 seconds on a 12GB GPU, flag for noisy, undesirable motifs, cluster the motifs, and provide other utilities to facilitate motif analyses. Using MotifCompendium, we consolidated the discovered motifs into a single, unified lexicon of 3,384 motifs that are predicted to drive TF binding and chromatin accessibility across cell contexts. From the lexicon, we observe that motifs from chromatin accessibility can be highly orthogonal to those from TF ChIP-seq, especially in different cell contexts (primary cells vs. cell lines). The lexicon also shows substantial overlap with existing TF binding motif databases, recapitulating, on average, ~93% of TF motifs from previous curated databases, while 10% of the lexicon represent entirely novel binding modes not captured by existing databases. We identify previously uncharacterized zinc finger-like motifs, composite motifs with two or more sub-motifs in preferential spacing, and cell-type-specific variants of motifs. This work provides a foundational resource of predictive models, a scalable computational framework for extracting sequence features, and the first unified lexicon of model-derived regulatory motifs.

Discussions

Related