How this index is built

Methodology

Every label in this index — evidence categories, body systems, symptoms, study types, safety signals — is generated by the rules described on this page. A tag records what published research talks about: research coverage, never a claim that a strain treats, cures, or prevents anything.

The corpus

Papers come from OpenAlex and PubMed, gathered per strain with quoted-phrase queries against strain names, tradenames, and culture-collection deposit numbers, so a paper enters the corpus because it names a strain we index — or, for a few single-species probiotics without a usable strain identifier, the species itself — not because it matched a broad topic. Papers are linked to strains at fetch and enrichment time, each link carrying a relevance level and a confidence score. Strain records themselves are curated from the published literature and culture-collection catalogs (ATCC, DSM, NCIMB, CNCM), and genomes come from NCBI.

This corpus is independent of the Winnow Atlas microplastics corpus; the two are built and maintained separately.

Genome pipeline

Each strain with a public genome runs through a six-stage pipeline. The output is each strain's per-category confidence, never a statement about what the strain does in a person. Paper matching is not a pipeline stage; strain–paper links are made when the corpus is fetched and enriched (see The corpus).

  1. Check availability. Confirm NCBI lists a public genome assembly for the strain. RefSeq assemblies are preferred over GenBank when both exist.
  2. Fetch assets. Download the assembly and its annotation from NCBI Datasets and link it to the strain by accession.
  3. Extract features. Scan the genome for genes and gene clusters relevant to the 10 evidence categories — including the three contaminant categories: microplastic, heavy metal, and mycotoxin & chemical interaction.
  4. Compute signature. Compute genome-level signals — a k-mer signature and surface-protein hydrophobicity — recorded as additional genome features.
  5. Roll up features. Aggregate extracted features into one rollup per category, each carrying an ordinal confidence. For the three contaminant categories, only host evidence — in vivo or gut-context research, and adhesion or adsorption gene families — can lift confidence to moderate or strong; environmental evidence (e.g., plastic or metal breakdown outside the body) is tracked separately and cannot lift confidence at all. Genome features extracted before the contaminant split carry no context yet and still count toward confidence until their genomes are re-extracted.
  6. Vectorize. Build the strain's feature and text vectors so similar strains can be found and mapped.

Evidence categories

Genome features and literature are organized into 10 categories. Each category's confidence for a strain reflects two inputs: how many relevant papers exist, and whatever the strain's genome screen contributes once one has run. No category is pinned or emphasized over the others.

Papers and genome features in a contaminant category are tagged host or environmental once they have been classified — host from in vivo or gut-context research, environmental from soil, water, landfill or bioreactor work. Re-classification of the existing corpus is still in progress, so many rows carry no tag yet.

For the three contaminant categories, only a host tag lifts a strain's confidence: a paper that is environmental, unclear, or not yet classified never counts toward a host claim. Genome features extracted before the split are the one exception — they carry no tag and are still counted, until their genomes are re-extracted against the current gene catalog.

Survival & Stability
Will it arrive alive — acid, bile, oxygen tolerance.
Gut Adhesion & Colonization
Sticks to the gut wall vs. passes through.
Microplastic Interaction
Binding & sequestration of microplastic particles.
Heavy Metal Interaction
Lead, cadmium, mercury — biosorption in the literature, resistance genes in the genome.
Mycotoxin & Chemical Interaction
Binding of mycotoxins, BPA and other plastic-derived chemicals.
Gut Barrier & Inflammation
SCFAs, tight-junction reinforcement, mucin support.
Immune Interaction
Immunomodulation, AMP production, T-reg signalling.
Metabolic Capabilities
Fiber fermentation, B-vitamins, GABA, histamine handling.
Safety & Transparency
What the genome screen looks for: AMR markers, toxin loci, assembly completeness.
Precision / Differentiation
Strain fingerprint, rare pathways, what makes it unique.

Confidence levels

Each category rollup carries an ordinal confidence — deliberately not a percentage, because genome evidence is qualitative. It reflects how much genomic signal was found, not how well the strain performs.

None
No genome-derived signal for this category, so its literature is all there is. A paper count is never gated by confidence; this level states only what the genome screen contributed.
Weak
A small number of relevant genes or partial pathways detected.
Moderate
Multiple relevant genes or complete pathways detected.
Strong
A robust, repeated genomic signal across the category's feature set.

Body systems

Papers are tagged with body systems by a conservative keyword scanner over each paper's title, abstract, and summary. Patterns aim to fire only on explicit mentions — bare "oral" is excluded because "oral administration" appears in most probiotics papers, bare "bone" would tag immunology papers via "bone marrow", and so on. Even so, keyword matches can still misfire on unrelated senses of a term; known false positives are corrected as they surface. One paper can carry several systems.

A tag means the paper's text names that system: research coverage, not measured effect. The twelve systems:

Gut
Stomach, small intestine, colon — the largest microbial habitat in the body.
Immune
Systemic and mucosal immunity, including the gut-associated lymphoid tissue.
Skin
Skin microbiome and barrier function.
Vaginal
Vaginal flora and urogenital health.
Oral
Mouth, tongue, and dental biofilms.
Urinary
Bladder and urinary tract.
Brain & Mood
Gut-brain axis, neuroactive metabolite production.
Metabolic
Glucose, lipids, and weight regulation.
Respiratory
Airways and lungs, including upper respiratory infections and allergy.
Liver
Hepatic health via the gut-liver axis, including fatty liver.
Cardiovascular
Heart and vessels — cholesterol, blood pressure, vascular health.
Bone & Joint
Bone density and joint health, including arthritis research.

Symptoms & conditions

Symptom and condition tags work the same way — a conservative scanner aimed at explicit mentions, with the same limits. Terms common in probiotics papers for unrelated reasons are excluded: bare "lactose" is a growth substrate in half the fermentation corpus, "ruminal bloat" is cattle husbandry, and livestock virus names like "porcine epidemic diarrhea" never count as diarrhea research.

A tag is a statement that the paper's text names the symptom — research coverage, never evidence of benefit. The twenty-four symptoms and conditions:

IBS
Irritable bowel syndrome — bloating, cramping, irregularity.
IBD
Inflammatory bowel disease — Crohn's and ulcerative colitis.
Lactose intolerance
Impaired digestion of lactose.
Constipation
Slow transit and infrequent bowel movements.
Diarrhea
Acute or chronic loose stools.
Eczema / Atopic dermatitis
Inflammatory skin condition often co-occurring with gut dysbiosis.
Allergic rhinitis
Seasonal or perennial nasal allergy.
Bacterial vaginosis
Vaginal microbial imbalance.
Urinary tract infection
Bacterial infection of the urinary system.
Anxiety
Mood and stress, often linked to gut-brain signalling.
Histamine intolerance
Difficulty metabolizing dietary or microbial histamine.
Bloating & gas
Abdominal distension, flatulence, and trapped gas.
Acne
Inflammatory skin breakouts, often studied via the gut-skin axis.
Depression
Low mood and depressive symptoms, studied via gut-brain signalling.
Sleep
Sleep quality, insomnia, and disturbed sleep patterns.
H. pylori
Helicobacter pylori colonization of the stomach lining.
Respiratory infections
Common colds, influenza, and upper respiratory tract infections.
Yeast infections
Candidiasis of the vaginal or oral microbiome.
Weight & obesity
Obesity, overweight, and weight-management research.
High cholesterol
Cholesterol and blood-lipid levels, including in-vitro lowering mechanisms.
Blood pressure
Hypertension research, including ACE-inhibitory fermented dairy.
Oral health
Dental caries, gum disease, and bad breath.
Food allergies
Food and milk allergies, often studied via early-life immune training.
Abdominal pain
Abdominal pain, cramping, and visceral discomfort.

Study types

Each paper is classified into one study type by an AI classifier and shown as a badge in research lists. Classification is automated and audited — papers labeled RCT are re-checked, and mis-classified ones are demoted — but individual labels can still be wrong; follow the link to the paper itself before leaning on one.

RCT
A randomized controlled trial in humans.
Meta-analysis
A quantitative synthesis of multiple trials.
Review
A narrative or systematic review of existing research.
Observational
A human study without randomization — cohort, case-control, or cross-sectional.
Animal
An in-vivo study in animals.
In-vitro
A laboratory study in cells or culture, outside a living organism.
WGS
A whole-genome sequencing or genomic-analysis paper.
Other
A paper the classifier could not place in the types above.

Relevance levels

Every strain–paper link carries a relevance level, extracted when the paper is matched to the strain, plus a numeric confidence that orders papers on strain pages.

Primary
The strain is a main subject of the paper.
Secondary
The strain is one of several organisms studied.
Mention
The strain is named in passing — context, comparison, or citation.

Safety screening

The signal at the top of a strain page reports coverage, not fitness for consumption. Green means the check could run and returned a result. When it could not run, because no public genome is linked, the strain page shows no signal at all rather than a cautionary one — the gap is in our records, not in the organism. Where that gap changes what a section can show, the section says so; where we have nothing to show either way, the page stays quiet rather than listing what it does not have.

Nothing in this index assesses whether a strain is safe to eat. The corpus is organisms that appear in contaminant-interaction literature, which includes species no one should consume, and we record no hazard, pathogenicity, or GRAS/QPS status against any of them. Read the taxonomy and the papers, not the dots.

Sequenced
Whether a public genome assembly exists for the strain. Without one, the genome-based screens below cannot run.

Taxonomy

Strains are grouped by genus and species using current NCBI taxonomy. In 2020 the genus Lactobacillus was split into 25 genera, so many familiar strains now carry new genus names — strain pages show the current name and note the former one ("was Lactobacillus"). Identity is anchored to NCBI taxon IDs and culture-collection deposit numbers wherever they exist. These are different systems and a strain page shows both: a deposit number such as ATCC 53103 identifies a physical culture held by a collection, while an NCBI taxon ID identifies a node in NCBI's taxonomy. A strain page names which rank its taxon ID has, because NCBI stopped creating strain-level nodes for bacteria years ago and most strains here can only be anchored to their species.

Isolation origin

Where a strain was first isolated, curated from the published literature and culture-collection catalogs. It describes the strain's provenance and nothing else — an environmental isolate is not thereby unsuitable, and a human isolate is not thereby beneficial.

Most of the directory is still Unknown, so strain cards omit the badge rather than stamping "we don't know" on every one; the filter can still select it.

Human
Isolated from a human source — gut, breast milk, or mucosa.
Food
Isolated from a fermented food, dairy, or another food matrix.
Environmental
Isolated from soil, water, plant material, or another environmental sample.
Other
A documented isolation source that fits none of the categories above.
Unknown
No isolation source recorded. This is the default and it covers most of the directory, so read it as a gap in our curation rather than as a fact about the strain.

Strain relations

The nodemap and the related-strain rail draw edges between strains. Curated relations are used when they exist; until an editorial pass records them, every edge is synthesized instead from the strains' published-summary embeddings — nearest neighbours by cosine distance. In practice that means most of what you see today is narrative similarity, which is a statement about how two strains are written about, not about their biology.

Solid
Genome similarity between two strains, from a curated relation.
Dashed
Narrative similarity: the two strains' published summaries sit close together in the same embedding space the related-strain rail uses.
Dotted
The two strains were studied together, from a curated relation.

A curated relation can also record that two strains appear in the same product, or that one antagonizes or complements the other. Those carry no distinct line style yet and draw as a plain edge, so don't read an undecorated line as one of the three above.