← Back to Ribosome Atlas

Ribosome Atlas Tutorial

Roma Nagle*, Yekaterina Shulgina*, Jamie H.D. Cate *equal contribution

Ribosome Atlas is an interactive resource for exploring ribosomal RNA (rRNA) diversity across microbial life. It combines large-scale phylogenies, multiple sequence alignments, and structural annotations so you can compare 16S and 23S rRNA sequence variation across any branch of the bacterial or archaeal tree of life. By linking evolutionary context with alignment- and structure-derived features, Ribosome Atlas is meant to support research in RNA evolution, function, and structure modeling.

Contents
  1. About this resource
  2. Dataset statistics
  3. Quick start
  4. Navigating the phylogenetic tree
  5. Specifying alignment positions
  6. Interpreting the visualizations
  7. Secondary structure annotations
  8. Alignment notation & symbols
  9. E. coli–based indexing
  10. How to cite

1. About This Resource

Ribosome Atlas is built on GTDB release r220 (April 2024), the Genome Taxonomy Database's standardized, phylogenetically consistent taxonomy for Bacteria and Archaea. For every genome in GTDB r220, we extracted the 16S and 23S ribosomal RNA gene sequences (where present in the assembly) and aligned them within their domain against the corresponding Rfam covariance model, then reindexed the alignment to E. coli numbering (see below) so that a given column means the same thing regardless of which organism you're looking at.

For any clade you select — from a single genus up to an entire domain — the site lets you:

Everything is computed live from the underlying alignments, so results always reflect exactly the clade and positions you specify.

2. Dataset Statistics

All figures below are computed directly from the current production databases (GTDB r220).

MetricBacteriaArchaea
Species-representative genomes107,2355,869
Phyla19722
Classes54267
Orders1,864170
Families4,897566
Genera23,1131,851
AlignmentLength (E. coli columns)Genomes with full-length coverage
16S rRNA1,542 positions107,235 / 107,235 bacteria; 5,869 / 5,869 archaea (>99.9%)
23S rRNA2,904 positions58,032 / 107,235 bacteria (54%); 3,686 / 5,869 archaea (63%)
23S coverage is lower than 16S because many GTDB assemblies are incomplete or fragmented at the (much longer) 23S locus. Genomes without a recovered 23S sequence are simply excluded from all 23S-derived statistics (base composition, entropy, deletion frequency, consensus) — they are not counted as gaps.

3. Quick Start

The fastest way to get a figure out of Ribosome Atlas:

  1. Pick Bacteria or Archaea from the Choose tree dropdown.
  2. Pick a taxonomic rank from Choose a level (or leave it at Domain to browse the whole tree), then pick a specific group in the Options box.
  3. Pick a view level to control how tree tips are colored/grouped (e.g. by genus or species).
  4. Type one or more positions into 16S positions and/or 23S positions, using E. coli numbering (e.g. 1492-1510).
  5. Click Generate.

The tree, alignment, base composition, entropy, deletion frequency, and consensus panels all update together. Clicking any node in the tree drills into that clade and regenerates every panel. Sections 4–7 below walk through each of these steps and panels in detail.

Step 1 – Choose a domain

Use the Choose tree dropdown to select either Bacteria or Archaea. This loads the corresponding GTDB phylogeny.

Step 2 – Choose a taxonomic level

Use the Choose a level dropdown to pick the rank you want to explore:

Step 3 – Select a group

The Options dropdown is populated with all groups at the chosen level. Select one to load its subtree in the tree panel on the left.

Step 4 – Choose a view level

The Choose a view level dropdown controls how the tree leaves are colored and grouped. For example, selecting Species colors each tip by species, while selecting Genus collapses colors at the genus level. This is independent of the level you used to filter.

The tree is interactive — clicking on a node will drill into that clade and update all panels. The breadcrumb below the controls shows which clade is currently displayed.

5. Specifying Alignment Positions

The two text inputs — 16S positions and 23S positions — let you select specific columns from the ribosomal RNA alignments to display. Positions are numbered relative to the E. coli reference (see E. coli–based indexing).

Input format

Example workflow

Enter 1492-1510 in the 16S field to examine the 3′ end of the small subunit rRNA across the selected clade, then click Generate.

After entering your positions, click Generate. The tree and alignment panels will update to reflect your selection. You can leave either field blank to skip that molecule.

Positions that fall outside the alignment or are only gaps will be silently skipped. If no valid positions remain, an error message will appear in red below the inputs.

6. Interpreting the Visualizations

Phylogenetic tree (left panel)

The tree shows the evolutionary relationships among organisms in the selected clade. Tips are labeled and colored by the view level you chose. Internal nodes can be clicked to zoom into a subtree.

Base composition plot

For each selected alignment position, the stacked bar chart shows the proportion of each nucleotide (A, U, G, C) present in the clade at that column of the alignment. Gap characters are excluded from all counts, so the chart reflects only organisms that have a nucleotide at that position.

Colors follow standard nucleotide conventions:

How to read the chart:

Each bar corresponds to one alignment position. When you specify multiple positions or a range, the bars are arranged left to right in the order you entered them. Gaps between non-contiguous ranges are shown as visual separators.

The base composition is computed from species-level alignments for all organisms in the selected clade, regardless of the view level chosen in the tree panel.

Position detail page

Clicking on any bar in the base composition chart pins a summary panel for that position. At the bottom of the panel, click Open full details to open a dedicated page in a new tab. That page contains:

This page is useful for identifying exactly which organisms contribute to a conserved or variable position, and for cross-referencing sequence variation with phylogenetic placement.

Hovering over a row in the pinned summary panel highlights which clades carry that nucleotide before you open the full details page.

Shannon entropy plot

Shannon entropy is a measure of sequence variability at each alignment position. It is computed from the same per-position nucleotide frequencies as the base composition chart (gaps excluded).

The y-axis is log-scaled to better distinguish low-entropy (highly conserved) positions. Positions with zero variance — where every organism has the same nucleotide — cannot be shown on a log scale and are marked with * at the base of the chart.

Deletion frequency plot

For each selected position, this bar chart shows the fraction of organisms in the clade that have a real deletion (an alignment gap, - or .) at that column, rather than a nucleotide. The y-axis is linear, 0–100%.

The denominator is organisms with either a base or a gap at that position — genomes with missing or fragmented sequence (~) covering that region are excluded from both the numerator and denominator, so a genome that simply wasn't sequenced through that region doesn't get counted as having a deletion.

This is a distinct signal from Shannon entropy: entropy describes variability among the bases that are present, while deletion frequency describes how often a base is present at all.

Domain Consensus

The Domain Consensus row shows the consensus sequence computed from all organisms in the selected domain (all Archaea or all Bacteria), at the positions you specified. It uses the same R/Y/N notation as the Clade Consensus (described below) but represents the full-domain background rather than the selected subtree.

Use this row to see whether a position is universally conserved across the domain or whether the pattern you observe in your selected clade is domain-wide or clade-specific.

Clade Consensus

The Clade Consensus row shows the consensus sequence for the specific clade you have selected (the organisms currently shown in the tree). It summarizes nucleotide identity at each position using the R/Y/N rules described in Alignment Notation & Symbols below.

Comparing the Domain Consensus and Clade Consensus side by side lets you quickly identify positions where your selected clade diverges from the broader domain pattern.

Alignment panel (right panels)

The alignment SVG shows the actual nucleotide sequence for each organism at the selected positions, arranged to match the tree on the left. This lets you directly compare sequence variation across the phylogeny at the positions you specified.

7. Secondary Structure Annotations

Below the position inputs, two bars — one for 16S, one for 23S — show where your selected positions fall within the known secondary structure of the rRNA.

Domain map

The colored bar shows the major structural domains (5′, central, 3′ major, 3′ minor for 16S; domains I–VI for 23S), sized proportionally to their length in E. coli numbering. Whichever domain(s) overlap your selected positions light up.

Dot-bracket notation

Below the domain map, each selected range is shown in WUSS dot-bracket notation — the same secondary-structure format used by Rfam and Infernal. It was generated by aligning E. coli's own 16S/23S rRNA sequence to the domain-specific Rfam covariance model (16S: RF00177 for bacteria, RF01959 for archaea; 23S: RF02541 for bacteria, RF02540 for archaea) and reading off the resulting base-pairing. See Alignment Notation & Symbols for the character key.

Hover over any character to see its exact position (and pairing partner, if any).

Bracket boundaries here are derived directly from Infernal's alignment output, so a single traditional helix (as numbered in the classic H1–H45 literature convention) may appear as several separate stems if it contains a bulge or internal loop.

8. Alignment Notation & Symbols

This section is a single reference for every symbol used across the site's sequences, consensus rows, and structure bars.

Nucleotide bases & ambiguity codes

Individual organism sequences use standard IUPAC codes (T is mapped to U throughout, since these are RNA molecules):

SymbolMeaning
A / U / G / CUnambiguous base call (colors match the base composition legend)
RA or G (purine)
YC or U (pyrimidine)
SG or C
WA or U
KG or U
MA or C
B / D / H / VNot A / not C / not G / not U, respectively
NUnknown/unresolved base in this genome's sequence (any of A/U/G/C)
- or .Alignment gap: this organism has a real deletion at this position relative to E. coli
~Missing/fragmented sequence: this region wasn't recovered for this genome. Excluded entirely from base composition, entropy, deletion frequency, and consensus calculations — not counted as a base or a gap
When positions are ambiguous (R, Y, S, W, K, M, B, D, H, V, N), base composition and entropy calculations split them fractionally across the bases they represent (e.g. an R contributes 0.5 to A and 0.5 to G) rather than discarding them.

Consensus rows (Domain Consensus / Clade Consensus)

The Domain Consensus and Clade Consensus rows do not reproduce any single genome's sequence — each position is a population-level summary computed across all organisms in the domain or clade, using these rules:

SymbolMeaning
A / U / G / C≥95% of non-gap bases at this position agree on this nucleotide
R≥70% of non-gap bases are purines (A or G), without reaching the 95% single-base threshold
Y≥70% of non-gap bases are pyrimidines (C or U), without reaching the 95% single-base threshold
NNo clear majority — the clade/domain does not agree on a base here
(blank)No consensus call at all — every organism in the clade/domain is deleted at this position, or no base here clears the posterior probability threshold. Left blank rather than marked, since an absent call is not the same as an agreed-upon deletion
Note the double meaning of N: in a raw sequence it means "this genome's base call is unresolved." In a consensus row it means "this population doesn't agree on one base here." The two are computed completely differently even though they use the same letter.
"Non-gap bases" above means non-gap bases that clear the posterior probability threshold set under Advanced settings. Both consensus rows are recomputed whenever you change it, from exactly the base calls the base composition chart is counting — so a position's consensus letter always follows the bars drawn above it. Loosening the threshold admits more, lower-confidence bases and can change a letter.

Secondary structure (dot-bracket / WUSS notation)

SymbolMeaning
.Unpaired position
< > { } [ ] ( )Base-paired position; matching bracket characters of the same type pair with each other. Different bracket types (e.g. <> vs {}) mark pseudoknots — base pairs that cross rather than nest, such as the 16S central pseudoknot
orange, dottedThis position is base-paired, but its partner falls outside your currently selected range

Base composition colors

These same four colors are used consistently for base composition bars, consensus-row letters, and sequence text throughout the site.

9. E. coli–Based Indexing

Every alignment position in Ribosome Atlas — in the position input boxes, the consensus rows, the entropy/base-composition/deletion-frequency plots, the domain maps, and the dot-bracket strings — is numbered according to E. coli rRNA numbering, the convention long used in the ribosome literature (e.g. "A2451" for the peptidyl transferase center in 23S rRNA, or "helix 44" landmarks in 16S rRNA).

Why E. coli numbering?

E. coli is itself one of the GTDB-represented bacterial genomes. Because every alignment in this resource is built against the same Rfam covariance model, E. coli's own aligned 16S and 23S sequences can be used as a stable ruler: alignment columns that are gaps in E. coli (i.e., insertions present only in other lineages) are dropped from the position numbering, so that surviving column N always corresponds to position N of E. coli's mature rRNA — position 1–1,542 for 16S rRNA, position 1–2,904 for 23S rRNA — no matter which organism's row you're reading.

What this means in practice

10. How to Cite

If you use Ribosome Atlas in your research, please cite:

Nagle R, Cate JHD, Shulgina Y. Ribosome Atlas [Internet]. Available from: https://ribosomeatlas.org

(A companion manuscript is in preparation — this citation will be updated with full publication details once available.)

Please also cite the underlying genome taxonomy this resource is built on:

Genome Taxonomy Database (GTDB), release r220. See the GTDB website for the current recommended citation.