← BioTransfer GEO Dataset Finder
GEO series

A long context RNA foundation model for predicting transcriptome architecture

GSE280041 Homo sapiens Expression profiling by high throughput sequencing 104 samples 2024/10/22 GPL24676GPL34382
Summary
Linking DNA sequence to genomic function remains one of the grand challenges in genetics and genomics. Here, we combine large-scale single-molecule transcriptome sequencing of diverse cancer cell lines with cutting-edge machine learning to build LoRNASH, an RNA foundation model that learns how the nucleotide sequence of unspliced pre-mRNA dictates transcriptome architecture—the relative abundances and molecular structures of mRNA isoforms. Owing to its use of the StripedHyena architecture, LoRNASH handles extremely long sequence inputs at base-pair resolution (~65 kilobase pairs), allowing for quantitative, zero-shot prediction of all aspects of transcriptome architecture, including isoform abundance, isoform structure, and the impact of DNA sequence variants on transcript structure and abundance. We anticipate that our public data release and the accompanying frontier model will accelerate many aspects of RNA biotechnology. More broadly, we envision the use of LoRNASH as a foundation for fine-tuning of any transcriptome-related downstream prediction task, including cell-type specific gene expression, splicing, and general RNA processing.
Download
NCBI GEO page ↗ {# Names what the click gives you. "Open in finder" meant nothing to a visitor who arrived from a search engine and has never seen the tool. #} Find more human RNA-seq datasets →
Similar datasets

Search all human RNA-seq datasets in GEO →

Share this dataset

Metadata from NCBI GEO, cached and refreshed periodically — the NCBI page above is authoritative. Downloads link straight to NCBI/ENA; nothing is proxied through BioTransfer.