Isoform-level transcriptome assembly and lncRNA discovery from lrRNA-seq
lr-RNAseq. lncRNA. Isoform-level. Nextflow.
Long non-coding RNAs (lncRNAs) are crucial regulators of gene expression, yet their accurate identification remains challenging due to their low expression, diverse structures, and inherent complexity of its isoforms. While long-read RNA sequencing (lrRNA-seq) has emerged as a powerful solution for resolving full-length transcripts and complex isoform architectures, the field lacks standardized, reproducible, and end-to-end computational frameworks specifically optimized for lncRNA discovery from raw sequencing data. We introduce LongNonCoder, a fully automated, scalable Nextflow pipeline designed for high-confidence, isoform-level lncRNA discovery and characterization, starting directly from raw lrRNA-seq reads. The workflow integrates sequential subworkflows for quality control, splice-aware alignment, transcriptome assembly and quantification, lncRNA identification, detailed structural characterization and data visualization. LongNonCoder allows several user-defined parameters and incorporates tools such as NanoComp, MultiQC, minimap2, Bambu, GffCompare, RNAmining, and biomaRt. A core feature is its dual-layer classification strategy, which evaluates novel isoforms based on structural features and coding potential independent of the biotype of its parent locus, ensuring high-confidence lncRNA assignment and resolving both constrained "mRNA-like" and extended non-canonical isoforms. We benchmarked LongNonCoder using public ONT and PacBio FLNC reads datasets, demonstrating high precision (>85%) and the capability to robustly assemble the transcriptome at the isoform level. The pipeline successfully isolated distinct lncRNA structural sub-types, including novel isoforms arising from known genes, as well as antisense, intergenic, and intron-retained isoforms, effectively separating them from canonical backgrounds. Furthermore, LongNonCoder's resource profiling highlights its scalability and efficient execution within a FAIR-compliant framework. Ultimately, by generating fully curated, downstream-ready structural and expression matrices, LongNonCoder standardizes the transition from raw lrRNA-seq data to deep architectural investigation of the long non-coding transcriptome, offering a reliable, reproducible, end-to-end, and flexible foundation for regulatory RNA research.