2017/06/26 by Huda Al-Nayyef, Christophe Guyeux, Al-Nayyef, Huda +3
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · #Annotation #Artificial intelligence #Bacterial Genetics and Biotechnology #Bacterial genome size #Biology #Computational biology #Computer science #Data mining #FOS: Biological sciences #Gene #Gene Annotation #Genetics #Genome #Genome project #Genomics (q-bio.GN) #Genomics and Phylogenetic Studies #Insertion sequence #Mobile genetic elements #Pipeline (software) #Probiotics and Fermented Foods #Programming language #Transposable element #Transposase #Transposition (logic) #q-bio.GN
paper · pdf · doi:10.48550/arxiv.1706.08267
arxiv created 2017/06/26 · openalex publication_date 2017/06/26 · arxiv updated 2017/06/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Insertion Sequences (ISs) are small DNA segments that have the ability of\nmoving themselves into genomes. These types of mobile genetic elements (MGEs)\nseem to play an essential role in genomes rearrangements and evolution of\nprokaryotic genomes, but the tools that deal with discovering ISs in an\nefficient and accurate way are still too few and not totally precise. Two main\nfactors have big effects on IS discovery, namely: genes annotation and\nfunctionality prediction. Indeed, some specific genes called "transposases" are\nenzymes that are responsible of the production and catalysis for such\ntransposition, but there is currently no fully accurate method that could\ndecide whether a given predicted gene is either a real transposase or not. This\nis why authors of this article aim at designing a novel pipeline for ISs\ndetection and classification, which embeds the most recently available tools\ndeveloped in this field of research, namely OASIS (Optimized Annotation System\nfor Insertion Sequence) and ISFinder database (an up-to-date and accurate\nrepository of known insertion sequences). As this latter depend on predicted\ncoding sequences, the proposed pipeline will encompass too various kinds of\nbacterial genes annotation tools (that is, Prokka, BASys, and Prodigal). A\ncomplete IS detection and classification pipeline is then proposed and tested\non a set of 23 complete genomes of Pseudomonas aeruginosa. This pipeline can\nalso be used as an investigator of annotation tools performance, which has led\nus to conclude that Prodigal is the best software for IS prediction. A deepen\nstudy regarding IS elements in P.aeruginosa has then been conducted, leading to\nthe conclusion that close genomes inside this species have also a close numbers\nof IS families and groups.\n