Research desk
Every gene, with its sources named
One page per gene in human, mouse or rat: what it does, the transcripts a design targets, drawn, where it is expressed, the protein it encodes and the company it keeps, the diseases, variants and constraint on record, the same gene in the other two species, the non-coding RNAs at its locus, and the papers. Each panel says where its data came from and when. Every page ends at the order door.
Worked examples
Six targets that show what a page holds
- TP53humana minus-strand gene with many isoforms and a MANE Select transcript
- Trp53mousethe mouse ortholog, on the plus strand, where MANE Select does not apply
- Tp53ratthe rat ortholog on the GRCr8 assembly
- BRCA1humana long gene whose exons are small next to its introns
- MALAT1humana long non-coding RNA, so no coding sequence to draw
- miR-21-5phumana mature microRNA, the 5p arm of the hairpin hsa-mir-21, with its other arm, the gene MIR21 and the same arm in mouse and rat
What a gene page holds
- Summary
- Transcripts
- Expression
- Protein
- Interactions
- Pathways
- Disease
- Variants
- Constraint
- Orthologs
- MicroRNAs
- Long non-coding RNAs
- Literature
Read from HGNC, NCBI Datasets, Ensembl, NCBI Gene, UniProt, GTEx, the Human Protein Atlas, InterPro, AlphaFold DB, PDBe, STRING, Reactome, Open Targets, ClinGen, ClinVar, gnomAD, the Alliance of Genome Resources, RGD, miRBase and Europe PMC, each named on its panel, in its own unit and release.
A microRNA page and a long non-coding RNA page hold the same kind of record for theirs.
Three more doors
Pathways, diseases and tissues
Each explorer lists what one named release of one source holds, in that source's own unit and order, and every gene on it opens the gene page above. Nothing is averaged across sources.
- By pathwayEvery pathway of one Reactome release in human, mouse or rat, with the genes Reactome places in it and the evidence code the mapping file writes.
- By diseaseEvery disease, phenotype and measurement of one Open Targets release, its ranked targets with the score named for what it is, and ClinGen's validity curations.
- By tissueHuman tissues from GTEx and the Human Protein Atlas, each ranking genes in its own unit, with the units printed and never mixed.
Glossary
Twelve words the pages use
Each opens in place. The same words appear on every gene page, meaning the same thing.
Isoform
One of the transcripts a gene produces by starting, ending or splicing differently. Isoforms of one gene share some exons and differ in others, which is why the map draws each as its own row: a sequence in a shared exon silences all of them, a sequence in an exon only one carries silences that one.
MANE Select
The one transcript per human gene that RefSeq and Ensembl have agreed is the reference, with the same sequence in both catalogues. It is the usual starting point for a design in human. Mouse and rat have no MANE set, so those pages label the RefSeq Select and the Ensembl canonical transcript instead, and the two need not be the same model.
Spliced length
The length of the mature transcript with the introns removed, which is the length of the RNA a knockdown oligonucleotide meets. It is far shorter than the gene's span on the chromosome, and it is the figure the tables show.
TPM
Transcripts per million: how many copies of a gene's RNA there would be among a million transcripts of a sample, after the length of the transcript is taken into account. It compares a gene's abundance across the tissues of one dataset; a TPM from another dataset or another method is not comparable to it. The expression panel shows the median TPM over the samples of each tissue, which is the middle sample, not the mean, so a gene expressed in a few cells of a mixed tissue reads low.
pLDDT
AlphaFold's confidence in its own model, given for every residue: how sure the model is of where that residue sits. A stretch of high confidence is typically a well-ordered domain; a stretch of low confidence is typically a flexible or disordered region rather than an error in the model. It is a statement about the prediction, not about what the protein does or whether the gene is a good target.
Post-translational modification
A change made to a protein after it is translated: a phosphate, an acetyl group or a ubiquitin added to a residue, a sugar or a lipid attached, a bond formed between two cysteines. UniProt records each at its residue from the literature, and the protein page draws them as pins along the chain. They mark where a protein is regulated, so a knockdown's effect on the protein can lag behind its effect on the RNA while the modified forms turn over; the RNA level is still the first thing to measure.
Resolution
For an experimental structure, how fine the detail is that the data resolve, in angstrom: a smaller figure is a sharper structure. It is a property of the X-ray or electron microscopy data, stated by the entry's depositors, and a solution NMR entry carries none because the method does not produce one. The protein page prints it as PDBe states it, and prints nothing where the entry has none.
LOEUF
gnomAD's loss-of-function observed over expected upper bound fraction. For each gene it compares how many loss-of-function variants were seen in the sequenced population with how many were expected, and reports the upper end of the confidence range. A low value means the population carries far fewer such variants than expected, so losing one copy is likely selected against; a high value means the gene tolerates them. It describes human genetics, not what a knockdown will do in an experiment.
Ortholog
The gene in another species that descends from the same ancestral gene, typically with the same role. Three databases vote on the match and the page shows all three rather than one merged answer, because they can disagree. An ortholog being the same gene does not mean one oligonucleotide sequence fits both species; that is checked against each transcript at design.
Strand
Which of the chromosome's two strands the gene is read from. On the minus strand the transcript runs from high coordinates to low, so exon 1 has the highest genomic position. The map always draws the transcript 5' to 3', left to right, so exon 1 is at the left whatever the strand, and the caption says which strand it is.
Gapmer
An antisense oligonucleotide built with a central stretch of DNA-like bases flanked on both sides by modified bases. Where it binds its target RNA, the central gap lets the cell's RNase H1 cut the RNA, and the flanks strengthen the binding and protect the oligonucleotide from nucleases. It is the design for knocking a transcript down, and the RNA level is the first thing to measure.
Steric block
An oligonucleotide that binds its target RNA and stays there without recruiting RNase H, so the transcript is not cut. What it stops is whatever would have happened at that site: a ribosome starting translation, a splice site being read, a microRNA or a protein binding. The RNA level does not fall, so the readout is the protein, the splice product or the downstream effect, not a knockdown by qPCR.
The gene pages are the first of the research desk. The microRNA pages are keyed on the mature name and the long non-coding RNA pages on species and symbol.
The tools