Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 24 additions & 5 deletions _fleet/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -187,11 +187,30 @@ ontologies, because they are how NaturalProductMech cites its corpus. MIBiG is a
seeded grounding source; NPAtlas is a cross-reference target only, since its
licence bars ingestion into a CC BY 4.0 corpus.

`prefix_census.py`'s prefix alternation is a hand-maintained list and is known to
be incomplete: TOGO, UTEX and CCAP are absent although comparable registries
(MediaDive, DSMZ, ATCC, GOLD) are present. TOGO is CultureMech's second-largest
structured namespace at 2,833 occurrences, so the heatmap currently understates
it. Adding a prefix changes the heatmap's columns, so it needs a full rescan.
`prefix_census.py`'s prefix alternation is a hand-maintained list. Its `norm`
table folds every namespace named for or qualified by a counted registry into
that registry: case and alternate names (`gold:` and `GOLD:`, `SwissProt:` and
`UniProt:`, `TAXON:` and `NCBITaxon:`, `CAS-RN:` and `CAS:`, `TC:` and `TCDB:`), and entity-qualified
namespaces (`kegg.compound:`, `mediadive.medium:`, `gtdb.genome:`,
`uniprot.location:`, `RHEA-COMP:`), as `gold.ecosystem` and `pubchem.compound`
always did (#271). Reference and curator collections named for a registry are
not its terms and stay out: `GO_REF:` and `PO_REF:` are literature-like (#84
lists GO_REF as a CITATION candidate), and `GOC:` is curator attribution, which
#84 proposes never to count (#282). Mechs mix lowercase bioregistry spellings with upper-case ones,
TaxonMech most heavily (#279). The list was measured by scanning every record at the
#120 pins with a pattern allowing dots, underscores and hyphens; a spelling a
Mech adopts later is missed until someone scans again. Until #84's fold (#244,
#255, #270) those spellings went uncounted, 1.49 million GOLD and 556,160 BacDive
identifiers in TaxonMech among them. `build_subsets.py` folds the same spellings
for heatmap columns. Rhea compounds, whose ids share values with Rhea reactions,
and PDB ligand codes, a different kind of identifier from entries, keep their own
term keys, so an overlap never pairs a compound with a reaction or shows a
ligand as an entry. Which
namespaces to count at all is still open on #84, which has the measured
inventory: CATH, CDD, PROSITE, StrainInfo, LPSN, TOGO, UNII and about a hundred
more are cited but not counted. Adding a prefix to the census does not add a
heatmap column (`build_data.VOC` decides those, and `PrefixListTests` keeps the
lists consistent), but it changes the vocabulary tile and needs a full rescan.

## Published-site refresh (September 24, 2026)

Expand Down
2 changes: 1 addition & 1 deletion _fleet/data/fleet_data.json

Large diffs are not rendered by default.

54 changes: 32 additions & 22 deletions _fleet/data/prefix_census.json
Original file line number Diff line number Diff line change
Expand Up @@ -59,13 +59,13 @@
"GO": 11688,
"DOI": 6636,
"NCBITaxon": 1438,
"UniProt": 993,
"PMID": 698,
"UniProt": 337,
"PDB": 277,
"PDB": 280,
"InterPro": 206,
"ECO": 145,
"SO": 69,
"CHEBI": 65,
"CHEBI": 66,
"Pfam": 22,
"ComplexPortal": 16,
"METPO": 14,
Expand All @@ -79,32 +79,33 @@
"ProteinTraitsMech": {
"files": 429293,
"prefixes": {
"InterPro": 1480044,
"UniProt": 658574,
"PMID": 657527,
"RHEA": 613456,
"InterPro": 1480046,
"UniProt": 659208,
"PMID": 657552,
"RHEA": 624625,
"Pfam": 521966,
"GO": 421586,
"ARO": 397717,
"CHEBI": 374424,
"NCBITaxon": 358320,
"CHEBI": 374596,
"NCBITaxon": 358325,
"RO": 278407,
"PDB": 235957,
"PDB": 250830,
"DOI": 193845,
"NCBIfam": 139525,
"EC": 119999,
"EC": 120213,
"BFO": 74488,
"ComplexPortal": 20753,
"SO": 19302,
"TCDB": 4705,
"TCDB": 4899,
"KEGG": 2018,
"CL": 260,
"KEGG": 245,
"METPO": 244,
"PATO": 124,
"PO": 53,
"UBERON": 43,
"MESH": 14,
"ECO": 4,
"MESH": 1
"PubChem": 1
}
},
"NaturalProductMech": {
Expand All @@ -114,9 +115,11 @@
"MIBiG": 10133,
"DOI": 8437,
"PMID": 5703,
"PubChem": 3290,
"NPAtlas": 1756,
"CHEBI": 1659,
"UniProt": 613,
"ChEMBL": 41,
"RHEA": 14
}
},
Expand All @@ -126,30 +129,32 @@
"CHEBI": 22104,
"ARO": 16653,
"CAS": 1823,
"KEGG": 1461,
"PMID": 1207,
"NCBITaxon": 1085,
"PubChem": 592,
"DrugBank": 508,
"ChEMBL": 415,
"PDB": 339,
"UniProt": 278,
"PHIPO": 217,
"DOI": 216,
"PDB": 6,
"GO": 3
}
},
"MediaIngredientMech": {
"files": 2953,
"prefixes": {
"CHEBI": 8220,
"CAS": 727,
"CHEBI": 8224,
"CAS": 1463,
"MediaDive": 538,
"MICRO": 351,
"MESH": 326,
"MESH": 328,
"NCIT": 306,
"FOODON": 273,
"ENVO": 110,
"KEGG": 61,
"PMID": 29,
"MediaDive": 24,
"UBERON": 23,
"DOI": 9,
"BTO": 3,
Expand All @@ -160,6 +165,8 @@
"files": 6288,
"prefixes": {
"CHEBI": 160678,
"MediaDive": 4843,
"KOMODO": 3262,
"KEGG": 1406,
"MICRO": 356,
"FOODON": 223,
Expand All @@ -168,7 +175,6 @@
"PMID": 68,
"NCBITaxon": 63,
"ENVO": 13,
"MediaDive": 10,
"DSMZ": 9,
"BacDive": 1,
"NCIT": 1,
Expand All @@ -179,13 +185,17 @@
"files": 625960,
"prefixes": {
"NCBITaxon": 9250364,
"GTDB": 648112,
"GOLD": 1492983,
"GTDB": 783276,
"BacDive": 556160,
"DOI": 402247,
"IMG": 68303,
"DSMZ": 31553,
"PMID": 20558,
"ATCC": 8870
}
},
"_as_of": "2026-09-24",
"_as_of": "2026-09-25",
"_revisions": {
"HabitatMech": "b16e3099478a7ea55b82142836e9a118ae062cf6",
"CommunityMech": "8505a56d644b02fe87f776be9d667db4cf3f5c5d",
Expand Down
57 changes: 30 additions & 27 deletions _fleet/data/subsets_summary.json
Original file line number Diff line number Diff line change
Expand Up @@ -251,10 +251,10 @@
]
},
"CommunityMech|ProteinTraitsMech": {
"n": 689,
"n": 690,
"by": {
"CHEBI": 276,
"NCBITaxon": 226,
"NCBITaxon": 227,
"GO": 186,
"UBERON": 1
},
Expand Down Expand Up @@ -316,9 +316,9 @@
]
},
"CommunityMech|MediaIngredientMech": {
"n": 240,
"n": 241,
"by": {
"CHEBI": 238,
"CHEBI": 239,
"ENVO": 2
},
"ex": [
Expand Down Expand Up @@ -533,11 +533,11 @@
]
},
"CellStructureMech|ProteinTraitsMech": {
"n": 1058,
"n": 1059,
"by": {
"GO": 737,
"UniProt": 155,
"NCBITaxon": 77,
"NCBITaxon": 78,
"InterPro": 64,
"CHEBI": 15,
"Pfam": 10
Expand Down Expand Up @@ -660,9 +660,9 @@
]
},
"ProteinTraitsMech|NaturalProductMech": {
"n": 794,
"n": 795,
"by": {
"NCBITaxon": 487,
"NCBITaxon": 488,
"UniProt": 156,
"CHEBI": 148,
"RHEA": 3
Expand All @@ -683,10 +683,11 @@
]
},
"ProteinTraitsMech|AntibioticMech": {
"n": 2435,
"n": 2575,
"by": {
"ARO": 1740,
"CHEBI": 594,
"PDB": 140,
"NCBITaxon": 79,
"UniProt": 20,
"GO": 2
Expand All @@ -707,9 +708,9 @@
]
},
"ProteinTraitsMech|MediaIngredientMech": {
"n": 798,
"n": 803,
"by": {
"CHEBI": 797,
"CHEBI": 802,
"UBERON": 1
},
"ex": [
Expand All @@ -728,9 +729,9 @@
]
},
"ProteinTraitsMech|CultureMech": {
"n": 233,
"n": 234,
"by": {
"CHEBI": 221,
"CHEBI": 222,
"NCBITaxon": 10,
"UBERON": 2
},
Expand All @@ -750,9 +751,9 @@
]
},
"ProteinTraitsMech|TaxonMech": {
"n": 7444,
"n": 7445,
"by": {
"NCBITaxon": 7444
"NCBITaxon": 7445
},
"ex": [
{
Expand Down Expand Up @@ -853,9 +854,10 @@
]
},
"AntibioticMech|MediaIngredientMech": {
"n": 324,
"n": 345,
"by": {
"CHEBI": 324
"CHEBI": 324,
"CAS": 21
},
"ex": [
{
Expand Down Expand Up @@ -986,22 +988,22 @@
"CellStructureMech|InterPro": 33,
"CellStructureMech|METPO": 5,
"CellStructureMech|NCBITaxon": 531,
"CellStructureMech|PDB": 19,
"CellStructureMech|PDB": 20,
"CellStructureMech|Pfam": 8,
"CellStructureMech|UniProt": 65,
"CellStructureMech|UniProt": 184,
"ProteinTraitsMech|ARO": 7452,
"ProteinTraitsMech|CHEBI": 30122,
"ProteinTraitsMech|CHEBI": 30290,
"ProteinTraitsMech|GO": 110744,
"ProteinTraitsMech|InterPro": 120090,
"ProteinTraitsMech|KEGG": 231,
"ProteinTraitsMech|KEGG": 1967,
"ProteinTraitsMech|METPO": 119,
"ProteinTraitsMech|NCBITaxon": 98697,
"ProteinTraitsMech|NCBITaxon": 98699,
"ProteinTraitsMech|PATO": 81,
"ProteinTraitsMech|PDB": 51718,
"ProteinTraitsMech|PDB": 52174,
"ProteinTraitsMech|Pfam": 107455,
"ProteinTraitsMech|RHEA": 29550,
"ProteinTraitsMech|UBERON": 43,
"ProteinTraitsMech|UniProt": 159595,
"ProteinTraitsMech|UniProt": 160190,
"NaturalProductMech|CHEBI": 405,
"NaturalProductMech|MIBiG": 3115,
"NaturalProductMech|NCBITaxon": 3115,
Expand All @@ -1012,11 +1014,12 @@
"AntibioticMech|CAS": 1797,
"AntibioticMech|CHEBI": 2814,
"AntibioticMech|GO": 3,
"AntibioticMech|KEGG": 1154,
"AntibioticMech|NCBITaxon": 163,
"AntibioticMech|PDB": 5,
"AntibioticMech|PDB": 331,
"AntibioticMech|UniProt": 47,
"MediaIngredientMech|BTO": 1,
"MediaIngredientMech|CAS": 287,
"MediaIngredientMech|CAS": 739,
"MediaIngredientMech|CHEBI": 2232,
"MediaIngredientMech|ENVO": 26,
"MediaIngredientMech|FOODON": 99,
Expand All @@ -1030,7 +1033,7 @@
"CultureMech|KEGG": 329,
"CultureMech|NCBITaxon": 30,
"CultureMech|UBERON": 72,
"TaxonMech|GTDB": 118801,
"TaxonMech|GTDB": 119248,
"TaxonMech|NCBITaxon": 625960
}
}
9 changes: 8 additions & 1 deletion _fleet/fleet_fragment.html
Original file line number Diff line number Diff line change
Expand Up @@ -824,9 +824,16 @@ <h2 id="shared-vocabulary">Shared vocabulary</h2>
if (more > refs.length) out += ' <em>+' + fmt(more - refs.length) + ' more</em>';
return out;
}
// Term keys that count toward another column (build_subsets TERM_SPACE): a
// column filter such as "PDB:" must still find PDB-CCD ligand codes (#284).
var TERM_COLUMN = { "PDB-CCD": "PDB", "RHEA-COMP": "RHEA" };
function termText(t) {
var column = TERM_COLUMN[t.id.split(":")[0]];
return (t.id + " " + t.l + (column ? " " + column + ":" : "")).toLowerCase();
}
function renderTermList(box, doc) {
var q = (box.querySelector("input") || {}).value || "", shown = parseInt(box.dataset.shown || "40", 10);
var terms = doc.terms.filter(function (t) { return !q || (t.id + " " + t.l).toLowerCase().indexOf(q.toLowerCase()) >= 0; });
var terms = doc.terms.filter(function (t) { return !q || termText(t).indexOf(q.toLowerCase()) >= 0; });
box.querySelector("ol").innerHTML = terms.slice(0, shown).map(function (t) {
return '<li><b>' + esc(t.l || t.id) + '</b> <code>' + esc(t.id) + '</code><div class="in"><span>' + esc(doc.a) + ' (' + fmt(t.na) + '):</span> ' + recLinks(doc.base[doc.a], t.a, t.na) + '</div><div class="in"><span>' + esc(doc.b) + ' (' + fmt(t.nb) + '):</span> ' + recLinks(doc.base[doc.b], t.b, t.nb) + '</div></li>';
}).join("") || '<li><em>No terms match.</em></li>';
Expand Down
1 change: 1 addition & 0 deletions assets/fleet/cells/AntibioticMech--KEGG.json

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion assets/fleet/cells/AntibioticMech--PDB.json

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion assets/fleet/cells/CellStructureMech--PDB.json
Original file line number Diff line number Diff line change
@@ -1 +1 @@
{"mech":"CellStructureMech","prefix":"PDB","base":"https://culturebotai.github.io/CellStructureMech/pages/structures/","total":19,"records":[["appendage/omce_cytochrome_nanowire.html","OmcE cytochrome nanowire"],["appendage/omcs_cytochrome_nanowire.html","OmcS cytochrome nanowire"],["appendage/omcz_cytochrome_nanowire.html","OmcZ cytochrome nanowire"],["cytoskeleton/mamk_filament.html","MamK filament"],["energy_complex/cytochrome_bd_i_ubiquinol_oxidase_complex.html","cytochrome bd-I ubiquinol oxidase complex"],["energy_complex/cytochrome_o_ubiquinol_oxidase_complex.html","cytochrome o ubiquinol oxidase complex"],["energy_complex/fumarate_reductase_complex.html","fumarate reductase complex"],["energy_complex/methyl_coenzyme_m_reductase_complex.html","methyl coenzyme M reductase complex"],["energy_complex/nitrogenase_complex.html","nitrogenase complex"],["energy_complex/proton_transporting_atp_synthase_complex.html","proton-transporting ATP synthase complex"],["energy_complex/respiratory_chain_complex_i.html","respiratory chain complex I"],["energy_complex/respiratory_chain_complex_ii_succinate_dehydrogenase.html","respiratory chain complex II (succinate dehydrogenase)"],["energy_complex/respiratory_chain_complex_iii.html","respiratory chain complex III"],["energy_complex/respiratory_chain_complex_iv.html","respiratory chain complex IV"],["energy_complex/sodium_translocating_nadh_quinone_reductase_complex.html","sodium-translocating NADH:quinone reductase complex"],["other/dna_gyrase_complex.html","DNA gyrase complex"],["other/exodeoxyribonuclease_v_complex.html","exodeoxyribonuclease V complex"],["other/single_stranded_dna_binding_protein_complex.html","single-stranded DNA-binding protein complex"],["secretion_system/esx_3_type_vii_secretion_system_membrane_complex.html","ESX-3 type VII secretion system membrane complex"]]}
{"mech":"CellStructureMech","prefix":"PDB","base":"https://culturebotai.github.io/CellStructureMech/pages/structures/","total":20,"records":[["appendage/omce_cytochrome_nanowire.html","OmcE cytochrome nanowire"],["appendage/omcs_cytochrome_nanowire.html","OmcS cytochrome nanowire"],["appendage/omcz_cytochrome_nanowire.html","OmcZ cytochrome nanowire"],["cytoskeleton/mamk_filament.html","MamK filament"],["energy_complex/cytochrome_bd_i_ubiquinol_oxidase_complex.html","cytochrome bd-I ubiquinol oxidase complex"],["energy_complex/cytochrome_o_ubiquinol_oxidase_complex.html","cytochrome o ubiquinol oxidase complex"],["energy_complex/fumarate_reductase_complex.html","fumarate reductase complex"],["energy_complex/methyl_coenzyme_m_reductase_complex.html","methyl coenzyme M reductase complex"],["energy_complex/nitrogenase_complex.html","nitrogenase complex"],["energy_complex/proton_transporting_atp_synthase_complex.html","proton-transporting ATP synthase complex"],["energy_complex/respiratory_chain_complex_i.html","respiratory chain complex I"],["energy_complex/respiratory_chain_complex_ii_succinate_dehydrogenase.html","respiratory chain complex II (succinate dehydrogenase)"],["energy_complex/respiratory_chain_complex_iii.html","respiratory chain complex III"],["energy_complex/respiratory_chain_complex_iv.html","respiratory chain complex IV"],["energy_complex/sodium_translocating_nadh_quinone_reductase_complex.html","sodium-translocating NADH:quinone reductase complex"],["other/dna_gyrase_complex.html","DNA gyrase complex"],["other/exodeoxyribonuclease_v_complex.html","exodeoxyribonuclease V complex"],["other/single_stranded_dna_binding_protein_complex.html","single-stranded DNA-binding protein complex"],["ribonucleoprotein/exosome_rnase_complex.html","exosome (RNase complex)"],["secretion_system/esx_3_type_vii_secretion_system_membrane_complex.html","ESX-3 type VII secretion system membrane complex"]]}
2 changes: 1 addition & 1 deletion assets/fleet/cells/CellStructureMech--UniProt.json

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion assets/fleet/cells/MediaIngredientMech--CAS.json

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion assets/fleet/cells/ProteinTraitsMech--CHEBI.json

Large diffs are not rendered by default.

Loading
Loading