Skip to content
All projects

FIG. 02.5 — Project notes

  • Independent
  • Finished

CARIOGENOME: Comparative Genomics of Cariogenic and Commensal Oral Streptococci

A pure-Python comparative-genomics pipeline that compares 6 virulence-associated genes with 8 housekeeping controls across 22 complete genomes from 5 oral streptococcal species, against six pre-registered predictions.

Exploratory research — no clinical claims

Associations between sequence patterns and gene categories. They do not show that any gene or variant causes caries, and the project makes no clinical, diagnostic or therapeutic claim.

Median dN/dS, virulence vs housekeeping
0.156 vs 0.019
Sliding windows with a CI above 1
0 of 579
Amino-acid identity, virulence vs housekeeping
99.06% vs 99.90%
Forest plot of Cliff's delta with 95% confidence intervals for every virulence-versus-housekeeping comparison in S. mutans; only conservation and dN/dS differ significantly.
FIG. 02.5 — Virulence vs housekeeping effect sizes · S. mutans · Figure from the project repository

01Problem

S. mutans is strongly associated with dental caries, while its relatives S. sanguinis, S. gordonii, S. mitis and S. salivarius are mostly commensal. The predictions were written in HYPOTHESIS.md before any analysis.

02Question

Do genes associated with cariogenicity differ from housekeeping genes in conservation, selection pressure or compositional signature?

03Data

22 complete NCBI RefSeq chromosomes: 10 S. mutans strains and 3 strains each of S. sanguinis, S. gordonii, S. mitis and S. salivarius. Panel: gtfB, gtfC, gtfD, spaP, ftf and luxS vs recA, rpoB, gyrB, gyrA, sodA, pheS, atpD and tuf, plus 16S rRNA. Structure: GtfC, PDB 3AIE (Ito et al. 2011, J Mol Biol 408:177, PMID 21354427).

04Methods

  • M1 Retrieval and QC: Bio.Entrez download of complete RefSeq chromosomes; orthologs by reciprocal best hit (k-mer prefilter plus Smith-Waterman, BLOSUM62) with a documented QC rule (284 of 286 records pass).
  • M2 Composition: GC, GC1–3, GC skew, RSCU and the Codon Adaptation Index; gene-level Mann-Whitney tests, Cliff's δ with bootstrap CIs and Benjamini-Hochberg correction.
  • M3 Conservation: pairwise global alignments, a center-star progressive MSA with codon back-translation, and Henikoff-weighted Shannon entropy per column.
  • M4 Phylogenetics: K2P distances, neighbor-joining and UPGMA with 100 bootstrap replicates, and a concatenated housekeeping reference tree; discordance by Robinson-Foulds distance and supported conflicting splits.
  • M5 Selection: Nei-Gojobori dN/dS pooled over strain pairs, a 1000-replicate codon bootstrap, a synonymous-saturation check, sliding windows and per-species replication, cross-checked against Bio.codonalign.
  • M6–M7 Motifs and structure: GH70 conserved motifs and the conservation rank of the catalytic residues; conservation mapped onto PDB 3AIE with a distance-to-active-site correlation.
  • Validation on simulated data: the true 22-taxon topology is recovered exactly (normalized RF = 0.0) and ω with Spearman ρ = 0.991 and a median relative error of 8.8%.

05Tools

  • Python
  • Biopython
  • NumPy
  • SciPy
  • pandas
  • matplotlib
  • py3Dmol
  • Streamlit
  • pytest

06Visualizations

Bar chart of dN/dS per gene within S. mutans with 95% confidence intervals, next to a comparison of the six virulence-associated and eight housekeeping genes.
FIG. 02.5.1Figure from the project repository ·Pooled NG86 dN/dS within S. mutans for each gene with 95% codon-bootstrap CIs, and the virulence-vs-control comparison, showing purifying selection in all genes but weaker constraint on virulence-associated genes.
Charts of within-species dN/dS for each gene in S. mutans and the commensal species, with virulence-associated homologs elevated in each.
FIG. 02.5.2Figure from the project repository ·Within-species dN/dS for every gene in every species with at least three strains, showing that virulence-gene homologs in commensal species also have elevated omega.
Map of GtfC residues colored by conservation beside a scatter plot of conservation against distance from the catalytic center.
FIG. 02.5.3Figure from the project repository ·GtfC (PDB 3AIE chain A) C-alpha atoms projected to 2D and colored by GH70 conservation, and conservation vs distance to the catalytic center (Spearman rho with 95% CI).
Dashboard overview tab with summary metrics and the effect-size forest plot.
FIG. 02.5.4Figure from the project repository ·The Streamlit dashboard’s overview tab, with summary metrics and the effect-size forest plot.

07Findings

  • Relaxed purifying selection, not positive selection: within S. mutans, median dN/dS is 0.156 for virulence-associated genes vs 0.019 for housekeeping genes (Cliff's δ = 0.88 [0.50, 1.00], BH q = 0.012). The difference comes from dN, not dS; every gene has dN/dS below 1, and none of 579 sliding windows has a CI above 1.
  • The pattern is shared with commensals: homologs in commensal species also have elevated dN/dS (e.g. S. gordonii median 0.138 vs 0.011) — an exploratory analysis added after seeing the data — so it most likely reflects secreted and surface-protein families rather than a cariogenicity-specific signature.
  • Prediction P5 narrowly failed: the GtfC catalytic residues D477, E515 and D588 are invariant across 11 GH70 enzymes, but 20.4% of alignment columns are also invariant, so they rank in the top 10.2% — just outside the pre-registered top 10%. The rule was not changed after the run.
  • Lower conservation: mean amino-acid identity among S. mutans strains is 99.06% vs 99.90% (δ = −0.96 [−1.00, −0.75], q = 0.0055).
  • No compositional signature (GC, GC3, CAI and GC skew: all q ≥ 0.33), no phylogenetic evidence of transfer between species, and conservation on the crystal structure falls with distance from the active site (Spearman ρ = −0.30 [−0.36, −0.23]).

08Limitations

  • It does not show a cariogenicity-specific signature: the elevated dN/dS is also seen in commensal homologs.
  • Small gene panel (6 vs 8 genes): only large class differences can be detected, and non-significant results mean “no evidence of a difference”.
  • Distance-based phylogenetics (NJ and UPGMA on K2P distances) and an approximate center-star alignment rather than MUSCLE or MAFFT.
  • Limited strain sampling (10 S. mutans, 3 per commensal); within-species ω reflects polymorphism, and between-species dN/dS is unusable because synonymous sites are saturated in 37 of 42 comparisons.
  • Orthology by reciprocal best hit can miss orthologs within paralog families, and the virulence set is heterogeneous (luxS is a metabolic enzyme found in every species).
  • Sequence signatures do not establish function: nothing here shows that any gene or variant causes disease, and the project makes no clinical, diagnostic or therapeutic claim.