How this index is built
Methodology
Every label in this index — evidence categories, body systems, symptoms, study types, safety signals — is generated by the rules described on this page. A tag records what published research talks about: research coverage, never a claim that a strain treats, cures, or prevents anything.
The corpus
Papers come from OpenAlex and PubMed, gathered per strain with quoted-phrase queries against strain names, tradenames, and culture-collection deposit numbers, so a paper enters the corpus because it names a strain we index — or, for a few single-species probiotics without a usable strain identifier, the species itself — not because it matched a broad topic. Papers are linked to strains at fetch and enrichment time, each link carrying a relevance level and a confidence score. Strain records themselves are curated from the published literature and culture-collection catalogs (ATCC, DSM, NCIMB, CNCM), and genomes come from NCBI.
This corpus is independent of the Winnow Atlas microplastics corpus; the two are built and maintained separately.
Genome pipeline
Each strain with a public genome runs through a six-stage pipeline. The output is each strain's per-category confidence, never a statement about what the strain does in a person. Paper matching is not a pipeline stage; strain–paper links are made when the corpus is fetched and enriched (see The corpus).
- Check availability. Confirm NCBI lists a public genome assembly for the strain. RefSeq assemblies are preferred over GenBank when both exist.
- Fetch assets. Download the assembly and its annotation from NCBI Datasets and link it to the strain by accession.
- Extract features. Scan the genome for genes and gene clusters relevant to the 10 evidence categories — including the three contaminant categories: microplastic, heavy metal, and mycotoxin & chemical interaction.
- Compute signature. Compute genome-level signals — a k-mer signature and surface-protein hydrophobicity — recorded as additional genome features.
- Roll up features. Aggregate extracted features into one rollup per category, each carrying an ordinal confidence. For the three contaminant categories, only host evidence — in vivo or gut-context research, and adhesion or adsorption gene families — can lift confidence to moderate or strong; environmental evidence (e.g., plastic or metal breakdown outside the body) is tracked separately and cannot lift confidence at all. Genome features extracted before the contaminant split carry no context yet and still count toward confidence until their genomes are re-extracted.
- Vectorize. Build the strain's feature and text vectors so similar strains can be found and mapped.
Evidence categories
Genome features and literature are organized into 10 categories. Each category's confidence for a strain reflects two inputs: how many relevant papers exist, and whatever the strain's genome screen contributes once one has run. No category is pinned or emphasized over the others.
Papers and genome features in a contaminant category are tagged host or environmental once they have been classified — host from in vivo or gut-context research, environmental from soil, water, landfill or bioreactor work. Re-classification of the existing corpus is still in progress, so many rows carry no tag yet.
For the three contaminant categories, only a host tag lifts a strain's confidence: a paper that is environmental, unclear, or not yet classified never counts toward a host claim. Genome features extracted before the split are the one exception — they carry no tag and are still counted, until their genomes are re-extracted against the current gene catalog.
- Survival & Stability
- Will it arrive alive — acid, bile, oxygen tolerance.
- Gut Adhesion & Colonization
- Sticks to the gut wall vs. passes through.
- Microplastic Interaction
- Binding & sequestration of microplastic particles.
- Heavy Metal Interaction
- Lead, cadmium, mercury — biosorption in the literature, resistance genes in the genome.
- Mycotoxin & Chemical Interaction
- Binding of mycotoxins, BPA and other plastic-derived chemicals.
- Gut Barrier & Inflammation
- SCFAs, tight-junction reinforcement, mucin support.
- Immune Interaction
- Immunomodulation, AMP production, T-reg signalling.
- Metabolic Capabilities
- Fiber fermentation, B-vitamins, GABA, histamine handling.
- Safety & Transparency
- What the genome screen looks for: AMR markers, toxin loci, assembly completeness.
- Precision / Differentiation
- Strain fingerprint, rare pathways, what makes it unique.
Confidence levels
Each category rollup carries an ordinal confidence — deliberately not a percentage, because genome evidence is qualitative. It reflects how much genomic signal was found, not how well the strain performs.
- None
- No genome-derived signal for this category, so its literature is all there is. A paper count is never gated by confidence; this level states only what the genome screen contributed.
- Weak
- A small number of relevant genes or partial pathways detected.
- Moderate
- Multiple relevant genes or complete pathways detected.
- Strong
- A robust, repeated genomic signal across the category's feature set.
Body systems
Papers are tagged with body systems by a conservative keyword scanner over each paper's title, abstract, and summary. Patterns aim to fire only on explicit mentions — bare "oral" is excluded because "oral administration" appears in most probiotics papers, bare "bone" would tag immunology papers via "bone marrow", and so on. Even so, keyword matches can still misfire on unrelated senses of a term; known false positives are corrected as they surface. One paper can carry several systems.
A tag means the paper's text names that system: research coverage, not measured effect. The twelve systems:
- Gut
- Stomach, small intestine, colon — the largest microbial habitat in the body.
- Immune
- Systemic and mucosal immunity, including the gut-associated lymphoid tissue.
- Skin
- Skin microbiome and barrier function.
- Vaginal
- Vaginal flora and urogenital health.
- Oral
- Mouth, tongue, and dental biofilms.
- Urinary
- Bladder and urinary tract.
- Brain & Mood
- Gut-brain axis, neuroactive metabolite production.
- Metabolic
- Glucose, lipids, and weight regulation.
- Respiratory
- Airways and lungs, including upper respiratory infections and allergy.
- Liver
- Hepatic health via the gut-liver axis, including fatty liver.
- Cardiovascular
- Heart and vessels — cholesterol, blood pressure, vascular health.
- Bone & Joint
- Bone density and joint health, including arthritis research.
Symptoms & conditions
Symptom and condition tags work the same way — a conservative scanner aimed at explicit mentions, with the same limits. Terms common in probiotics papers for unrelated reasons are excluded: bare "lactose" is a growth substrate in half the fermentation corpus, "ruminal bloat" is cattle husbandry, and livestock virus names like "porcine epidemic diarrhea" never count as diarrhea research.
A tag is a statement that the paper's text names the symptom — research coverage, never evidence of benefit. The twenty-four symptoms and conditions:
- IBS
- Irritable bowel syndrome — bloating, cramping, irregularity.
- IBD
- Inflammatory bowel disease — Crohn's and ulcerative colitis.
- Lactose intolerance
- Impaired digestion of lactose.
- Constipation
- Slow transit and infrequent bowel movements.
- Diarrhea
- Acute or chronic loose stools.
- Eczema / Atopic dermatitis
- Inflammatory skin condition often co-occurring with gut dysbiosis.
- Allergic rhinitis
- Seasonal or perennial nasal allergy.
- Bacterial vaginosis
- Vaginal microbial imbalance.
- Urinary tract infection
- Bacterial infection of the urinary system.
- Anxiety
- Mood and stress, often linked to gut-brain signalling.
- Histamine intolerance
- Difficulty metabolizing dietary or microbial histamine.
- Bloating & gas
- Abdominal distension, flatulence, and trapped gas.
- Acne
- Inflammatory skin breakouts, often studied via the gut-skin axis.
- Depression
- Low mood and depressive symptoms, studied via gut-brain signalling.
- Sleep
- Sleep quality, insomnia, and disturbed sleep patterns.
- H. pylori
- Helicobacter pylori colonization of the stomach lining.
- Respiratory infections
- Common colds, influenza, and upper respiratory tract infections.
- Yeast infections
- Candidiasis of the vaginal or oral microbiome.
- Weight & obesity
- Obesity, overweight, and weight-management research.
- High cholesterol
- Cholesterol and blood-lipid levels, including in-vitro lowering mechanisms.
- Blood pressure
- Hypertension research, including ACE-inhibitory fermented dairy.
- Oral health
- Dental caries, gum disease, and bad breath.
- Food allergies
- Food and milk allergies, often studied via early-life immune training.
- Abdominal pain
- Abdominal pain, cramping, and visceral discomfort.
Study types
Each paper is classified into one study type by an AI classifier and shown as a badge in research lists. Classification is automated and audited — papers labeled RCT are re-checked, and mis-classified ones are demoted — but individual labels can still be wrong; follow the link to the paper itself before leaning on one.
- RCT
- A randomized controlled trial in humans.
- Meta-analysis
- A quantitative synthesis of multiple trials.
- Review
- A narrative or systematic review of existing research.
- Observational
- A human study without randomization — cohort, case-control, or cross-sectional.
- Animal
- An in-vivo study in animals.
- In-vitro
- A laboratory study in cells or culture, outside a living organism.
- WGS
- A whole-genome sequencing or genomic-analysis paper.
- Other
- A paper the classifier could not place in the types above.
Relevance levels
Every strain–paper link carries a relevance level, extracted when the paper is matched to the strain, plus a numeric confidence that orders papers on strain pages.
- Primary
- The strain is a main subject of the paper.
- Secondary
- The strain is one of several organisms studied.
- Mention
- The strain is named in passing — context, comparison, or citation.
Safety screening
The signal at the top of a strain page reports coverage, not fitness for consumption. Green means the check could run and returned a result. When it could not run, because no public genome is linked, the strain page shows no signal at all rather than a cautionary one — the gap is in our records, not in the organism. Where that gap changes what a section can show, the section says so; where we have nothing to show either way, the page stays quiet rather than listing what it does not have.
Nothing in this index assesses whether a strain is safe to eat. The corpus is organisms that appear in contaminant-interaction literature, which includes species no one should consume, and we record no hazard, pathogenicity, or GRAS/QPS status against any of them. Read the taxonomy and the papers, not the dots.
- Sequenced
- Whether a public genome assembly exists for the strain. Without one, the genome-based screens below cannot run.
Taxonomy
Strains are grouped by genus and species using current NCBI taxonomy. In 2020 the genus Lactobacillus was split into 25 genera, so many familiar strains now carry new genus names — strain pages show the current name and note the former one ("was Lactobacillus"). Identity is anchored to NCBI taxon IDs and culture-collection deposit numbers wherever they exist. These are different systems and a strain page shows both: a deposit number such as ATCC 53103 identifies a physical culture held by a collection, while an NCBI taxon ID identifies a node in NCBI's taxonomy. A strain page names which rank its taxon ID has, because NCBI stopped creating strain-level nodes for bacteria years ago and most strains here can only be anchored to their species.
Isolation origin
Where a strain was first isolated, curated from the published literature and culture-collection catalogs. It describes the strain's provenance and nothing else — an environmental isolate is not thereby unsuitable, and a human isolate is not thereby beneficial.
Most of the directory is still Unknown, so strain cards omit the badge rather than stamping "we don't know" on every one; the filter can still select it.
- Human
- Isolated from a human source — gut, breast milk, or mucosa.
- Food
- Isolated from a fermented food, dairy, or another food matrix.
- Environmental
- Isolated from soil, water, plant material, or another environmental sample.
- Other
- A documented isolation source that fits none of the categories above.
- Unknown
- No isolation source recorded. This is the default and it covers most of the directory, so read it as a gap in our curation rather than as a fact about the strain.
Strain relations
The nodemap and the related-strain rail draw edges between strains. Curated relations are used when they exist; until an editorial pass records them, every edge is synthesized instead from the strains' published-summary embeddings — nearest neighbours by cosine distance. In practice that means most of what you see today is narrative similarity, which is a statement about how two strains are written about, not about their biology.
- Solid
- Genome similarity between two strains, from a curated relation.
- Dashed
- Narrative similarity: the two strains' published summaries sit close together in the same embedding space the related-strain rail uses.
- Dotted
- The two strains were studied together, from a curated relation.
A curated relation can also record that two strains appear in the same product, or that one antagonizes or complements the other. Those carry no distinct line style yet and draw as a plain edge, so don't read an undecorated line as one of the three above.