The rows and the map are the transcripts NCBI's RefSeq annotation and Ensembl's release list for the gene, each catalogue in its own group with its own labels (MANE Select, RefSeq Select, Ensembl canonical) and its own release and assembly named; a transcript's biotype is printed as the catalogue writes it, and every exon drawn is the catalogue's own placement. Where the two catalogues disagree, both are shown and neither is preferred.
A biotype is the class the Ensembl/GENCODE annotators assign to a gene, and to each of its transcripts, from the evidence for it. [R05] [R07] [R10] Ensembl's own documentation defines a biotype as A gene or transcript classification
[R60] and the lncRNA biotype as A non-coding gene/transcript >200bp in length
[R60], and defines antisense, sense intronic, sense overlapping and lincRNA each as a class of transcripts [R60]; ‘lncRNA’ is one of the broad gene classes beside protein-coding, pseudogene and small non-coding RNA [R05], and GENCODE defines a gene's biotype by its transcripts': the ‘biotype’ of the gene, i.e. the functional classification we set, is defined by the biotype of the transcripts it contains
[R58]. The card prints the gene's biotype as the release its data carries writes it. The biotypes are GENCODE's, describing biological function at the transcript level
[R07]; a transcript Ensembl labels retained_intron contains, as exon, sequence that is intronic in the locus's other transcripts, containing sequence that is intronic in other transcripts from the locus
[R05], or in the current paper's words models that contain retained introns (i.e. introns that have not been spliced out)
[R58]; the label is Ensembl's, and the card lists such models as Ensembl lists them.
MANE Select is the one transcript per gene that RefSeq and Ensembl/GENCODE jointly chose as representative, with exon sequences identical in both, so that either identifier can be used; the set was built for protein-coding genes and, from MANE v1.4, also carries a small number of long non-coding RNAs. The set identifies a representative transcript for each human protein-coding gene
[R09], and Ensembl added MANE Select transcripts for a small set of clinically important long non-coding RNAs
[R10], first in MANE v1.4 available in Ensembl release 113
[R10]; the label marks the transcript recommended for those situations where only a single transcript is needed
[R58]. [R11]
An NR accession is a curated RefSeq record of a non-protein-coding transcript, created by curators for a transcript whose exon structure they are confident of; an XR accession is a model produced by NCBI's annotation pipeline for the same class. RefSeq curators take a conservative approach to representing lncRNA genes, only manually creating RefSeqs (with a NR_ accession prefix) for high quality transcripts for which we have some certainty of the exon structure.
[R17] The XR records are model RefSeqs (XM_, XR_ and XP_ accessions, Table 1)
[R17] from the annotation pipeline.
- [R05] Frankish A, Diekhans M, Ferreira AM, Johnson R, Jungreis I, Loveland J, et al. (2019). GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Research 47:D766-D773. PMID 30357393, doi 10.1093/nar/gky955.
- [R07] Frankish A, Carbonell-Sala S, Diekhans M, Jungreis I, Loveland JE, Mudge JM, et al. (2023). GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Research 51:D942-D949. PMID 36420896, doi 10.1093/nar/gkac1071.
- [R09] Morales J, Pujar S, Loveland JE, Astashyn A, Bennett R, Berry A, et al. (2022). A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature 604:310-315. PMID 35388217, doi 10.1038/s41586-022-04558-8.
- [R10] Dyer SC, Austine-Orimoloye O, Azov AG, Barba M, Barnes I, Barrera-Enriquez VP, et al. (2025). Ensembl 2025. Nucleic Acids Research 53:D948-D957. PMID 39656687, doi 10.1093/nar/gkae1071.
- [R11] Yates AD, Austine-Orimoloye O, Azov AG, Barba M, Barnes I, Barrera-Enriquez VP, et al. (2026). Ensembl 2026. Nucleic Acids Research 54:D1053-D1060. .