2025/05/26 by Xingyu Chen, Shihao Ma, Chen, Xingyu +7 · 1 voice · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Artificial Intelligence (cs.AI) #DNA and Biological Computing #Evolutionary Algorithms and Applications #FOS: Biological sciences #FOS: Computer and information sciences #Genomics (q-bio.GN) #Machine Learning (cs.LG) #Modular Robots and Swarm Intelligence #cs.AI #cs.LG #q-bio.GN
paper · pdf · doi:10.48550/arxiv.2505.20578
openalex publication_date 2025/05/26 · arxiv published 2025/05/26 · arxiv updated 2025/05/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Designing regulatory DNA sequences that achieve precise cell-type-specific gene expression is crucial for advancements in synthetic biology, gene therapy and precision medicine. Although transformer-based language models (LMs) can effectively capture patterns in regulatory DNA, their generative approaches often struggle to produce novel sequences with reliable cell-specific activity. Here, we introduce Ctrl-DNA, a novel constrained reinforcement learning (RL) framework tailored for designing regulatory DNA sequences with controllable cell-type specificity. By formulating regulatory sequence design as a biologically informed constrained optimization problem, we apply RL to autoregressive genomic LMs, enabling the models to iteratively refine sequences that maximize regulatory activity in targeted cell types while constraining off-target effects. Our evaluation on human promoters and enhancers demonstrates that Ctrl-DNA consistently outperforms existing generative and RL-based approaches, generating high-fitness regulatory sequences and achieving state-of-the-art cell-type specificity. Moreover, Ctrl-DNA-generated sequences capture key cell-type-specific transcription factor binding sites (TFBS), short DNA motifs recognized by regulatory proteins that control gene expression, demonstrating the biological plausibility of the generated sequences.