Banca de QUALIFICAÇÃO: MARIA BÁRBARA BORGES DE SANTANA

Uma banca de QUALIFICAÇÃO de MESTRADO foi cadastrada pelo programa.
STUDENT : MARIA BÁRBARA BORGES DE SANTANA
DATE: 30/05/2026
TIME: 09:00
LOCAL: Google Meet
TITLE:

Isoform-level transcriptome assembly and lncRNA discovery from lrRNA-seq


KEY WORDS:

lr-RNAseq. lncRNA. Isoform-level. Nextflow.


PAGES: 52
BIG AREA: Ciências Biológicas
AREA: Biologia Geral
SUMMARY:

Long non-coding RNAs (lncRNAs) are crucial regulators of gene expression, yet their accurate identification remains challenging due to their low expression, diverse structures, and inherent complexity of its isoforms. While long-read RNA sequencing (lrRNA-seq) has emerged as a powerful solution for resolving full-length transcripts and complex isoform architectures, the field lacks standardized, reproducible, and end-to-end computational frameworks specifically optimized for lncRNA discovery from raw sequencing data. We introduce LongNonCoder, a fully automated, scalable Nextflow pipeline designed for high-confidence, isoform-level lncRNA discovery and characterization, starting directly from raw lrRNA-seq reads. The workflow integrates sequential subworkflows for quality control, splice-aware alignment, transcriptome assembly and quantification, lncRNA identification, detailed structural characterization and data visualization. LongNonCoder allows several user-defined parameters and incorporates tools such as NanoComp, MultiQC, minimap2, Bambu, GffCompare, RNAmining, and biomaRt. A core feature is its dual-layer classification strategy, which evaluates novel isoforms based on structural features and coding potential independent of the biotype of its parent locus, ensuring high-confidence lncRNA assignment and resolving both constrained "mRNA-like" and extended non-canonical isoforms. We benchmarked LongNonCoder using public ONT and PacBio FLNC reads datasets, demonstrating high precision (>85%) and the capability to robustly assemble the transcriptome at the isoform level. The pipeline successfully isolated distinct lncRNA structural sub-types, including novel isoforms arising from known genes, as well as antisense, intergenic, and intron-retained isoforms, effectively separating them from canonical backgrounds. Furthermore, LongNonCoder's resource profiling highlights its scalability and efficient execution within a FAIR-compliant framework. Ultimately, by generating fully curated, downstream-ready structural and expression matrices, LongNonCoder standardizes the transition from raw lrRNA-seq data to deep architectural investigation of the long non-coding transcriptome, offering a reliable, reproducible, end-to-end, and flexible foundation for regulatory RNA research.


COMMITTEE MEMBERS:
Externo à Instituição - GLORIA REGINA FRANCO - UFMG
Interno - 1507794 - RODRIGO JULIANI SIQUEIRA DALMOLIN
Interna - ***.117.554-** - THAIS GAUDENCIO DO REGO - UFPB
Presidente - ***.739.204-** - VINICIUS RAMOS HENRIQUES MARACAJA COUTINHO - USP
Notícia cadastrada em: 20/05/2026 12:10
SIGAA | Superintendência de Tecnologia da Informação - (84) 3342 2210 | Copyright © 2006-2026 - UFRN - sigaa12-producao.info.ufrn.br.sigaa12-producao