Ribosome Atlas is an interactive resource for exploring ribosomal RNA (rRNA) diversity across microbial life. It combines large-scale phylogenies, multiple sequence alignments, and structural annotations so you can compare 16S and 23S rRNA sequence variation across any branch of the bacterial or archaeal tree of life. By linking evolutionary context with alignment- and structure-derived features, Ribosome Atlas is meant to support research in RNA evolution, function, and structure modeling.
Ribosome Atlas is built on GTDB release r220 (April 2024), the Genome Taxonomy Database's standardized, phylogenetically consistent taxonomy for Bacteria and Archaea. For every genome in GTDB r220, we extracted the 16S and 23S ribosomal RNA gene sequences (where present in the assembly) and aligned them within their domain against the corresponding Rfam covariance model, then reindexed the alignment to E. coli numbering (see below) so that a given column means the same thing regardless of which organism you're looking at.
For any clade you select — from a single genus up to an entire domain — the site lets you:
Everything is computed live from the underlying alignments, so results always reflect exactly the clade and positions you specify.
All figures below are computed directly from the current production databases (GTDB r220).
| Metric | Bacteria | Archaea |
|---|---|---|
| Species-representative genomes | 107,235 | 5,869 |
| Phyla | 197 | 22 |
| Classes | 542 | 67 |
| Orders | 1,864 | 170 |
| Families | 4,897 | 566 |
| Genera | 23,113 | 1,851 |
| Alignment | Length (E. coli columns) | Genomes with full-length coverage |
|---|---|---|
| 16S rRNA | 1,542 positions | 107,235 / 107,235 bacteria; 5,869 / 5,869 archaea (>99.9%) |
| 23S rRNA | 2,904 positions | 58,032 / 107,235 bacteria (54%); 3,686 / 5,869 archaea (63%) |
The fastest way to get a figure out of Ribosome Atlas:
1492-1510).The tree, alignment, base composition, entropy, deletion frequency, and consensus panels all update together. Clicking any node in the tree drills into that clade and regenerates every panel. Sections 4–7 below walk through each of these steps and panels in detail.
Use the Choose tree dropdown to select either Bacteria or Archaea. This loads the corresponding GTDB phylogeny.
Use the Choose a level dropdown to pick the rank you want to explore:
The Options dropdown is populated with all groups at the chosen level. Select one to load its subtree in the tree panel on the left.
The Choose a view level dropdown controls how the tree leaves are colored and grouped. For example, selecting Species colors each tip by species, while selecting Genus collapses colors at the genus level. This is independent of the level you used to filter.
The two text inputs — 16S positions and 23S positions — let you select specific columns from the ribosomal RNA alignments to display. Positions are numbered relative to the E. coli reference (see E. coli–based indexing).
530530-5405, 10-20, 25Enter 1492-1510 in the 16S field to examine the 3′ end of the small subunit rRNA across the selected clade, then click Generate.
After entering your positions, click Generate. The tree and alignment panels will update to reflect your selection. You can leave either field blank to skip that molecule.
The tree shows the evolutionary relationships among organisms in the selected clade. Tips are labeled and colored by the view level you chose. Internal nodes can be clicked to zoom into a subtree.
For each selected alignment position, the stacked bar chart shows the proportion of each nucleotide (A, U, G, C) present in the clade at that column of the alignment. Gap characters are excluded from all counts, so the chart reflects only organisms that have a nucleotide at that position.
Colors follow standard nucleotide conventions:
How to read the chart:
Each bar corresponds to one alignment position. When you specify multiple positions or a range, the bars are arranged left to right in the order you entered them. Gaps between non-contiguous ranges are shown as visual separators.
Clicking on any bar in the base composition chart pins a summary panel for that position. At the bottom of the panel, click Open full details to open a dedicated page in a new tab. That page contains:
This page is useful for identifying exactly which organisms contribute to a conserved or variable position, and for cross-referencing sequence variation with phylogenetic placement.
Shannon entropy is a measure of sequence variability at each alignment position. It is computed from the same per-position nucleotide frequencies as the base composition chart (gaps excluded).
The y-axis is log-scaled to better distinguish low-entropy (highly conserved) positions. Positions with zero variance — where every organism has the same nucleotide — cannot be shown on a log scale and are marked with * at the base of the chart.
For each selected position, this bar chart shows the fraction of organisms in the clade that have a real deletion (an alignment gap, - or .) at that column, rather than a nucleotide. The y-axis is linear, 0–100%.
The denominator is organisms with either a base or a gap at that position — genomes with missing or fragmented sequence (~) covering that region are excluded from both the numerator and denominator, so a genome that simply wasn't sequenced through that region doesn't get counted as having a deletion.
This is a distinct signal from Shannon entropy: entropy describes variability among the bases that are present, while deletion frequency describes how often a base is present at all.
The Domain Consensus row shows the consensus sequence computed from all organisms in the selected domain (all Archaea or all Bacteria), at the positions you specified. It uses the same R/Y/N notation as the Clade Consensus (described below) but represents the full-domain background rather than the selected subtree.
Use this row to see whether a position is universally conserved across the domain or whether the pattern you observe in your selected clade is domain-wide or clade-specific.
The Clade Consensus row shows the consensus sequence for the specific clade you have selected (the organisms currently shown in the tree). It summarizes nucleotide identity at each position using the R/Y/N rules described in Alignment Notation & Symbols below.
Comparing the Domain Consensus and Clade Consensus side by side lets you quickly identify positions where your selected clade diverges from the broader domain pattern.
The alignment SVG shows the actual nucleotide sequence for each organism at the selected positions, arranged to match the tree on the left. This lets you directly compare sequence variation across the phylogeny at the positions you specified.
Below the position inputs, two bars — one for 16S, one for 23S — show where your selected positions fall within the known secondary structure of the rRNA.
The colored bar shows the major structural domains (5′, central, 3′ major, 3′ minor for 16S; domains I–VI for 23S), sized proportionally to their length in E. coli numbering. Whichever domain(s) overlap your selected positions light up.
Below the domain map, each selected range is shown in WUSS dot-bracket notation — the same secondary-structure format used by Rfam and Infernal. It was generated by aligning E. coli's own 16S/23S rRNA sequence to the domain-specific Rfam covariance model (16S: RF00177 for bacteria, RF01959 for archaea; 23S: RF02541 for bacteria, RF02540 for archaea) and reading off the resulting base-pairing. See Alignment Notation & Symbols for the character key.
Hover over any character to see its exact position (and pairing partner, if any).
This section is a single reference for every symbol used across the site's sequences, consensus rows, and structure bars.
Individual organism sequences use standard IUPAC codes (T is mapped to U throughout, since these are RNA molecules):
| Symbol | Meaning |
|---|---|
| A / U / G / C | Unambiguous base call (colors match the base composition legend) |
| R | A or G (purine) |
| Y | C or U (pyrimidine) |
| S | G or C |
| W | A or U |
| K | G or U |
| M | A or C |
| B / D / H / V | Not A / not C / not G / not U, respectively |
| N | Unknown/unresolved base in this genome's sequence (any of A/U/G/C) |
| - or . | Alignment gap: this organism has a real deletion at this position relative to E. coli |
| ~ | Missing/fragmented sequence: this region wasn't recovered for this genome. Excluded entirely from base composition, entropy, deletion frequency, and consensus calculations — not counted as a base or a gap |
The Domain Consensus and Clade Consensus rows do not reproduce any single genome's sequence — each position is a population-level summary computed across all organisms in the domain or clade, using these rules:
| Symbol | Meaning |
|---|---|
| A / U / G / C | ≥95% of non-gap bases at this position agree on this nucleotide |
| R | ≥70% of non-gap bases are purines (A or G), without reaching the 95% single-base threshold |
| Y | ≥70% of non-gap bases are pyrimidines (C or U), without reaching the 95% single-base threshold |
| N | No clear majority — the clade/domain does not agree on a base here |
| (blank) | No consensus call at all — every organism in the clade/domain is deleted at this position, or no base here clears the posterior probability threshold. Left blank rather than marked, since an absent call is not the same as an agreed-upon deletion |
| Symbol | Meaning |
|---|---|
| . | Unpaired position |
| < > { } [ ] ( ) | Base-paired position; matching bracket characters of the same type pair with each other. Different bracket types (e.g. <> vs {}) mark pseudoknots — base pairs that cross rather than nest, such as the 16S central pseudoknot |
| orange, dotted | This position is base-paired, but its partner falls outside your currently selected range |
These same four colors are used consistently for base composition bars, consensus-row letters, and sequence text throughout the site.
Every alignment position in Ribosome Atlas — in the position input boxes, the consensus rows, the entropy/base-composition/deletion-frequency plots, the domain maps, and the dot-bracket strings — is numbered according to E. coli rRNA numbering, the convention long used in the ribosome literature (e.g. "A2451" for the peptidyl transferase center in 23S rRNA, or "helix 44" landmarks in 16S rRNA).
E. coli is itself one of the GTDB-represented bacterial genomes. Because every alignment in this resource is built against the same Rfam covariance model, E. coli's own aligned 16S and 23S sequences can be used as a stable ruler: alignment columns that are gaps in E. coli (i.e., insertions present only in other lineages) are dropped from the position numbering, so that surviving column N always corresponds to position N of E. coli's mature rRNA — position 1–1,542 for 16S rRNA, position 1–2,904 for 23S rRNA — no matter which organism's row you're reading.
1492) refers to the same structural/functional site in every organism's alignment row, letting you compare that exact site across the whole tree of life.-/.) at that position — see notation.If you use Ribosome Atlas in your research, please cite:
Nagle R, Cate JHD, Shulgina Y. Ribosome Atlas [Internet]. Available from: https://ribosomeatlas.org
(A companion manuscript is in preparation — this citation will be updated with full publication details once available.)
Please also cite the underlying genome taxonomy this resource is built on:
Genome Taxonomy Database (GTDB), release r220. See the GTDB website for the current recommended citation.