annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing

dc.contributor.affiliationUniversidad Catolica de la Santisima Concepcion
dc.contributor.affiliationUniversidad de Concepcion
dc.contributor.affiliationUniversity of Manitoba
dc.contributor.affiliationUniversidad de Las Americas - Chile
dc.contributor.affiliationUniversidad Bernardo O'Higgins
dc.contributor.authorFarkas, Carlos
dc.contributor.authorRecabal, Antonia
dc.contributor.authorMella, Andy
dc.contributor.authorCandia-Herrera, Daniel
dc.contributor.authorGonzález Olivero, Maryori
dc.contributor.authorHaigh, Jody Jonathan
dc.contributor.authorTarifeño-Saldivia, Estefania
dc.contributor.authorCaprile, Teresa
dc.date.accessioned2024-09-03T19:21:25Z
dc.date.available2024-09-03T19:21:25Z
dc.date.issued2022
dc.description.abstractAbstract Background The advancement of hybrid sequencing technologies is increasingly expanding genome assemblies that are often annotated using hybrid sequencing transcriptomics, leading to improved genome characterization and the identification of novel genes and isoforms in a wide variety of organisms. Results We developed an easy-to-use genome-guided transcriptome annotation pipeline that uses assembled transcripts from hybrid sequencing data as input and distinguishes between coding and long non-coding RNAs by integration of several bioinformatic approaches, including gene reconciliation with previous annotations in GTF format. We demonstrated the efficiency of this approach by correctly assembling and annotating all exons from the chicken SCO-spondin gene (containing more than 105 exons), including the identification of missing genes in the chicken reference annotations by homology assignments. Conclusions Our method helps to improve the current transcriptome annotation of the chicken brain. Our pipeline, implemented on Anaconda/Nextflow and Docker is an easy-to-use package that can be applied to a broad range of species, tissues, and research areas helping to improve and reconcile current annotations. The code and datasets are publicly available at https://github.com/cfarkas/annotate_my_genomes
dc.description.sponsorshipFondo Nacional de Desarrollo Cientifico y Tecnologico, FONDECYT [1191860]; FONDECYT de Iniciacion [11190401]; CIHR; CancerCare Manitoba Foundation; This work was supported by Fondo Nacional de Desarrollo Cientifico y Tecnologico, FONDECYT [1191860 to T.C] and FONDECYT de Iniciacion [11190401 to ETS]. CF and JJH received partial funding from the CIHR and CancerCare Manitoba Foundation.
dc.format.mimetypeapplication/pdf
dc.identifier.citationGigaScience, 11, giac099. https://doi.org/10.1093/gigascience/giac099
dc.identifier.doihttps://doi.org/10.1093/gigascience/giac099
dc.identifier.folio11190401
dc.identifier.folio1191860
dc.identifier.issn2047-217X
dc.identifier.orcidhttps://orcid.org/0000-0002-6245-2622
dc.identifier.orcidhttps://orcid.org/0000-0003-2883-4739
dc.identifier.orcidhttps://orcid.org/0000-0003-3186-9617
dc.identifier.orcidhttps://orcid.org/0000-0002-7319-8922
dc.identifier.orcidhttps://orcid.org/0000-0001-5311-2661
dc.identifier.orcidhttps://orcid.org/0000-0002-0897-7049
dc.identifier.orcidhttps://orcid.org/0000-0001-8143-8482
dc.identifier.orcidhttps://orcid.org/0000-0001-8934-8756
dc.identifier.orcidhttps://orcid.org/0000-0001-6632-8875
dc.identifier.pmid36472574
dc.identifier.researcheridE-8251-2013
dc.identifier.researcheridAAD-4649-2019
dc.identifier.researcheridHGE-9982-2022
dc.identifier.rorhttps://ror.org/03y6k2j68
dc.identifier.rorhttps://ror.org/0460jpj73
dc.identifier.rorhttps://ror.org/00x0xhn70
dc.identifier.rorhttps://ror.org/005cmms77
dc.identifier.rorhttps://ror.org/05qer3y84
dc.identifier.rorhttps://ror.org/02gfys938
dc.identifier.rorhttps://ror.org/0166e9x11
dc.identifier.scopusauthorid55847466100
dc.identifier.scopusauthorid56711977600
dc.identifier.scopusauthorid57191289716
dc.identifier.scopusauthorid57998817400
dc.identifier.scopusauthorid57284745100
dc.identifier.scopusauthorid7103171838
dc.identifier.scopusauthorid56145387200
dc.identifier.scopusauthorid6507062024
dc.identifier.urihttps://repositorio.udla.cl/handle/udla/1635
dc.language.isoeng
dc.publisherOxford University Press (OUP)
dc.relation.fundingCancerCare Manitoba Foundation, CCMF
dc.relation.fundingCanadian Institutes of Health Research, IRSC
dc.relation.fundingFondo Nacional de Desarrollo Científico y Tecnológico, FONDECYT, (11190401, 1191860)
dc.relation.fundingFondo Nacional de Desarrollo Científico y Tecnológico, FONDECYT
dc.relation.fundingFondo Nacional de Desarrollo Cientifico y Tecnologico, FONDECYT [1191860]
dc.relation.fundingFONDECYT de Iniciacion [11190401]
dc.relation.fundingCIHR
dc.relation.fundingCancerCare Manitoba Foundation
dc.relation.isindexedbyWeb of Science
dc.rightsCreative Commons Attribution 4.0 International
dc.rights.accessrightsinfo:eu-repo/semantics/openAccess
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.sourceGIGASCIENCE
dc.source.urihttps://academic.oup.com/gigascience/article-abstract/doi/10.1093/gigascience/giac099/6874526
dc.subjectTranscriptome annotation
dc.subjectGenome Annotation pipeline
dc.subjectSCO-spondin
dc.subjecthybrid sequencing
dc.subject.oecd11 Ciencias Naturales
dc.subject.oecd21.6 Ciencias Biológicas
dc.titleannotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing
dc.title.alternativeannotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
dc.typejournal article
dc.type.coarhttp://purl.org/coar/resource_type/c_6501
dc.type.driverinfo:eu-repo/semantics/article
dc.udla.catalogadorCBM
oaire.citation.titleGIGASCIENCE
oaire.citation.volume11
oaire.fundingReference.awardNumber11190401
oaire.fundingReference.awardNumber1191860
oaire.fundingReference.funderNameAgencia Nacional de Investigación y Desarrollo (ANID)
udla.curacion.controljmvg
udla.oecd.area1 Ciencias Naturales
udla.oecd.subarea1.6 Ciencias Biológicas

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
510.pdf
Size:
5.12 MB
Format:
Adobe Portable Document Format

Collections