2024/05/09 by Daanish Mahajan, Chirag Jain, Mahajan, Daanish +3
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · Computer Science · #Chromosomal and Genetic Variations #DNA and Biological Computing #Evolutionary Algorithms and Applications #FOS: Biological sciences #FOS: Computer and information sciences #Genomics (q-bio.GN) #Information Theory (cs.IT)
paper · doi:10.48550/arxiv.2405.05734
openalex publication_date 2024/05/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
The repeat content and heterozygosity rate of a target genome are important factors in determining the feasibility of achieving a complete telomere-to-telomere assembly. The mathematical relationship between the required coverage and read length for the purpose of unique reconstruction remains unexplored for diploid genomes. We investigate the information-theoretic conditions that the given set of sequencing reads must satisfy to achieve the complete reconstruction of the true sequence of a diploid genome. We also analyze the standard greedy and de-Bruijn graph-based assembly algorithms. Our results show that the coverage and read length requirements of the assembly algorithms are considerably higher than the lower bound because both algorithms require the double repeats in the genome to be bridged. Finally, we derive the necessary conditions for the overlap graph-based assembly paradigm.