Oral-microbiome research. The signal in your mouth is real. It just doesn't travel well.

✓ reproduces preprintRead the preprint (PDF)4 unseen populations tested

What is this?

Can a spit test predict Type 2 Diabetes from the bacteria in your mouth? A model trained on some populations is tested on new populations it never saw. The score is an AUC: 0.5 is a coin flip, 1.0 is perfect. A point estimate below 0.5 points the wrong way; the confidence interval tells us whether that reversal is distinguishable from chance. The headline: transfer varies sharply across unseen cohorts, and in South Africa it is reliably reversed.

The investigation

four experiments, one thread · tap a step to jump

Tested across the world

trained here tested (color = AUC)
NHANES (US)China0.698Portugal0.621India/Mumbai0.564Mexico0.320South Africa

The model learned mainly from US and Chinese cohorts. Geography shows where each test occurred; it is not a measured population-distance axis. Transfer varies by cohort and reverses in South Africa.

Cross-population transfer

trained on pooled cohorts, scored on each unseen population
backwards0.30.40.50.60.70.80.9chancePortugalEuropean · saliva0.860.698above chanceIndia/MumbaiSouth Asian · plaque0.640.621uncertain vs chanceMexicoLatin American · saliva0.570.564uncertain vs chanceSouth AfricaAfrican · plaque0.850.320reversed
0.5 coin flip
1.0 perfecttransferwithin-cohortwhisker = transfer 95% CI
best: Portugal 0.698 · worst: South Africa 0.320

Cohorts

CohortPopulationSpecimenSamplesWithin-cohortTransfer AUC (95% CI)
PortugalEuropeansaliva50 (25/25)0.8590.698[0.544, 0.837]above chance
India/MumbaiSouth Asianplaque60 (40/20)0.6380.621[0.455, 0.782]uncertain vs chance
MexicoLatin Americansaliva87 (41/46)0.5660.564[0.441, 0.684]uncertain vs chance
South AfricaAfricanplaque88 (57/31)0.8530.320[0.200, 0.448]reversed

Within-cohort = signal learnable inside the cohort (5-fold CV). Transfer = the pooled model scored on it (never trained on it). The two estimates come from different fits: together they show learnable local structure that the pooled ranking does not preserve. For South Africa, transfer is below chance (0.320, 95% CI 0.200–0.448) despite strong within-cohort CV (0.853).

Which genera generalize?

experiment #2 · per-genus T2D direction across six populations

If the model reads South Africa backwards, the natural question is which oral genera keep a consistent diabetes direction across populations and which flip. The answer is stark: direction, not strength, is what fails to travel.

71%
of genera informative in 2+ populations reverse sign
15
genera hold one direction across 5+ populations (all weak)
56
sign-reversing flippers, Atopobium among them
US
n=9683
CN
n=624
PT
n=50
IN
n=60
MX
n=87
ZA
n=88
Direction-stable · 12 of 15
Peptoniphilus
Enterococcus
Alloscardovia
Lactobacillus
Streptococcus
Bifidobacterium
Fretibacterium
Johnsonella
Cloacibacterium
Treponema
Delftia
Schlegelella
Sign-reversing flippers · sharpest 14 of 56
Serratia
Faecalibacterium
Campylobacter
Lysinibacillus
Microbacterium
Burkholderia
Bacteroides
Ureaplasma
Paenibacillus
Actinomyces
Parabacteroides
Atopobium
Dermabacter
Eubacterium
control ← r → caserow = genus · column = population · value = point-biserial r with T2D
How many association-top markers flip, per cohort

Of the 20 genera with the strongest pooled training associations (a different ranking from model importance: Atopobium is not among them), the fraction that flip sign in each cohort. The relation to transfer is loose, not monotone: India and South Africa flip the same fraction at very different transfer AUCs.

Portugal
10%
markers flip · transfer 0.698
India/Mumbai
40%
markers flip · transfer 0.621
Mexico
50%
markers flip · transfer 0.564
South Africa
40%
markers flip · transfer 0.320
Test: a portability-first model (experiment #3)
stable features chosen leave-one-cohort-out (no leakage)

If direction is the problem, restrict the model to the direction-stable genera. It trades peak accuracy for safety: Portugal loses ground, but South Africa stops being read backwards and no cohort reverses.

chancePortugal0.5890.698India/Mumbai0.6160.621Mexico0.6000.564South Africa0.5150.320no longer reversed
reversed cohorts 1 → 0mean transfer 0.551 → 0.580drop Atopobium alone: South Africa 0.320 → 0.410 (partial, so the reversal is broad)

Point-biserial correlation between each sample's per-genus rank and its T2D label, computed within each cohort (leakage-free). NHANES magnitudes are diffuse (n=9683) so a cohort casts only a sign vote once |r| clears 0.05; the 71% reversal figure excludes NHANES. The portability test selects stable genera leave-one-cohort-out, so a cohort never informs its own feature set; a de-inversion toward 0.5 is a loss of confident-wrongness, not new skill. Associations are hypothesis-generating, not validated biomarkers: per-cohort samples are modest and several flippers are low-abundance or environmental taxa.

Population biology, or study effect?

experiment #8 · the honest reframe

Is the South Africa reversal African biology or a study artifact? The training pool holds four Chinese studies, same population, different labs. If directions flip among them too, batch alone can reverse a marker. They do: same-population studies flip almost as often as different populations.

61%
of markers flip among four same-population (Chinese) studies
71%
flip across distinct populations, only ~10 points more
0.163 vs -0.001
direction agreement: within-China vs cross-population
Atopobium, the headline marker, flips sign even between two Chinese studies
lower in diabeticshigher in diabeticsChinese sub-studies (one population)anhui yang+0.19shandong sun-0.46shanghai t2dm-0.14sichuan liu-0.25Distinct populationsNHANES+0.02China-0.04Portugal-0.07India/Mumbai+0.15Mexico-0.19South Africa+0.47

Per-genus point-biserial T2D association within each group. Same-population studies do agree a little more than different populations (0.163 vs -0.001 mean correlation), so population is not nothing, but most of the direction instability is present within a single population. The parsimonious reading: the reversal is not distinguishable from generic study-level instability, and a population-specific biological explanation is not supported by public data. Chinese sub-studies are small (n=45 to 347), so their per-genus signs are themselves noisy, which is precisely the point.

What makes a marker portable?

experiment #6 · genus prevalence & abundance vs stability

Are flippers just the rare taxa whose association sign is noisy? Partly. Direction-stable markers are about 6x more prevalent than flippers, but prevalence is no guarantee: a set of common genera reverse anyway.

1e-41e-31%1e-1100%prevalence (fraction of samples detected, log)mean abundance (log)Faecalibacterium: prevalence 12.9%, flipperFaecalibacteriumCellulosimicrobium: prevalence 0.4%, flipperDelftia: prevalence 1.7%, robust_downMogibacterium: prevalence 90.2%, weakAlloscardovia: prevalence 28.4%, robust_upBordetella: prevalence 0.3%, flipperOlsenella: prevalence 31.5%, weakHaematobacter: prevalence 2.1%, flipperDialister: prevalence 85.4%, weakPorphyromonas: prevalence 95.5%, weakCentipeda: prevalence 8.4%, weakMobiluncus: prevalence 14.2%, flipperLeptotrichia: prevalence 96.3%, weakAtopobium: prevalence 87.2%, flipperAtopobiumBacillus: prevalence 0.6%, weakParvimonas: prevalence 90.9%, weakAnoxybacillus: prevalence 0.6%, weakSelenomonas: prevalence 69.2%, weakDesulfobulbus: prevalence 23.7%, weakStreptococcus: prevalence 98.8%, robust_upStreptococcusShuttleworthia: prevalence 39.2%, weakRothia: prevalence 98.2%, weakFinegoldia: prevalence 3.6%, weakScardovia: prevalence 62.2%, weakAfipia: prevalence 1.0%, flipperCatonella: prevalence 81.5%, weakMicrococcus: prevalence 3.5%, weakDermabacter: prevalence 0.1%, flipperCapnocytophaga: prevalence 92.1%, flipperOdoribacter: prevalence 7.4%, weakAnaeroglobus: prevalence 56.9%, weakTreponema: prevalence 6.3%, robust_downNovosphingobium: prevalence 2.0%, weakMesorhizobium: prevalence 1.4%, weakProteus: prevalence 1.0%, weakRalstonia: prevalence 22.3%, weakSlackia: prevalence 14.1%, weakBifidobacterium: prevalence 52.5%, robust_upPaenibacillus: prevalence 1.0%, flipperEubacterium: prevalence 0.5%, flipperAlloprevotella: prevalence 95.6%, flipperAggregatibacter: prevalence 68.8%, flipperSolobacterium: prevalence 92.3%, weakRhodobacter: prevalence 2.7%, flipperVictivallis: prevalence 0.1%, flipperYersinia: prevalence 5.8%, flipperCampylobacter: prevalence 95.3%, flipperCorynebacterium: prevalence 92.5%, weakJonquetella: prevalence 0.5%, flipperBacteroides: prevalence 36.4%, flipperPeptococcus: prevalence 67.5%, flipperBurkholderia: prevalence 4.3%, flipperEikenella: prevalence 32.9%, flipperSphingomonas: prevalence 5.9%, weakSchlegelella: prevalence 0.4%, robust_downMycobacterium: prevalence 3.9%, weakCardiobacterium: prevalence 68.3%, flipperPedobacter: prevalence 0.3%, flipperHelicobacter: prevalence 1.3%, flipperMoraxella: prevalence 19.2%, weakMitsuokella: prevalence 2.4%, flipperOttowia: prevalence 17.2%, weakAnaerococcus: prevalence 6.3%, weakDolosigranulum: prevalence 1.1%, weakSimonsiella: prevalence 7.5%, weakOribacterium: prevalence 91.0%, flipperRoseomonas: prevalence 0.4%, flipperVariovorax: prevalence 0.3%, weakPseudoramibacter: prevalence 22.6%, robust_upBulleidia: prevalence 38.8%, weakSneathia: prevalence 6.7%, weakPeptoniphilus: prevalence 11.8%, robust_upBrucella: prevalence 2.8%, flipperLysinibacillus: prevalence 0.4%, flipperMicrobacterium: prevalence 1.1%, flipperUreaplasma: prevalence 0.5%, flipperEggerthella: prevalence 1.4%, flipperAchromobacter: prevalence 0.5%, flipperAlistipes: prevalence 6.0%, flipperEnhydrobacter: prevalence 0.3%, weakParascardovia: prevalence 18.6%, weakCryptobacterium: prevalence 34.2%, weakFusobacterium: prevalence 97.2%, flipperBrevundimonas: prevalence 2.7%, robust_downBradyrhizobium: prevalence 5.2%, weakSerratia: prevalence 1.2%, flipperAcinetobacter: prevalence 21.0%, robust_upDesulfovibrio: prevalence 9.2%, weakMoryella: prevalence 11.7%, flipperErythrobacter: prevalence 0.8%, weakBartonella: prevalence 0.1%, flipperDesulfomicrobium: prevalence 6.2%, weakAerococcus: prevalence 0.6%, weakGardnerella: prevalence 5.3%, weakComamonas: prevalence 0.8%, weakFilifactor: prevalence 56.3%, weakTannerella: prevalence 77.4%, weakGemella: prevalence 97.5%, weakVeillonella: prevalence 98.2%, flipperActinomyces: prevalence 97.7%, flipperParabacteroides: prevalence 5.9%, flipperLautropia: prevalence 71.9%, flipperStenotrophomonas: prevalence 14.1%, weakKingella: prevalence 70.6%, weakKocuria: prevalence 2.1%, weakAcidovorax: prevalence 1.1%, weakSegetibacter: prevalence 0.1%, flipperLactococcus: prevalence 6.3%, weakPeptostreptococcus: prevalence 86.5%, flipperPseudomonas: prevalence 20.6%, weakJeotgalicoccus: prevalence 0.3%, flipperAquamicrobium: prevalence 0.6%, flipperStaphylococcus: prevalence 20.7%, flipperPhocaeicola: prevalence 31.1%, flipperCupriavidus: prevalence 0.9%, weakFretibacterium: prevalence 54.6%, robust_downArsenicicoccus: prevalence 0.1%, flipperCloacibacterium: prevalence 16.6%, robust_downAkkermansia: prevalence 10.3%, flipperStomatobaculum: prevalence 88.4%, weakArcanobacterium: prevalence 0.1%, flipperNeisseria: prevalence 93.1%, flipperGranulicatella: prevalence 97.0%, flipperPrevotella: prevalence 97.2%, weakMegasphaera: prevalence 89.4%, weakLactobacillus: prevalence 51.2%, robust_upEnterococcus: prevalence 2.9%, robust_upHaemophilus: prevalence 96.9%, weakJanibacter: prevalence 0.8%, weakDietzia: prevalence 0.1%, weakButyrivibrio: prevalence 5.1%, flipperCaulobacter: prevalence 0.5%, weakLachnoanaerobaculum: prevalence 89.4%, weakPropionibacterium: prevalence 6.8%, weakKytococcus: prevalence 0.1%, flipperJohnsonella: prevalence 38.0%, robust_downPyramidobacter: prevalence 6.3%, weakEggerthia: prevalence 24.3%, weakAbiotrophia: prevalence 71.2%, weakdirection-stableflipperweak

Prevalence and abundance are label-free properties from the full raw count table. Stable markers median prevalence 21% vs flippers 3.6% (Mann-Whitney p=0.1282, a suggestive but underpowered trend with only 15 stable genera). The prevalent, abundant genera that reverse anyway (Faecalibacterium, Veillonella, Akkermansia, Actinomyces, and more) are the biologically interesting reversers, not sparsity artifacts. Takeaway for panel design: prefer prevalent, abundant genera, but test their direction per population.

Does going coarse help? No.

experiment #7 · community diversity vs specific taxa

Maybe a coarse summary, community diversity, transfers where specific taxa don't. It doesn't: diversity transfers worse and reverses in more cohorts. The direction instability is baked in at every level, down to a single number.

chancePortugal0.432India/Mumbai0.544Mexico0.451South Africa0.188taxa (139 genera)diversity (3 features)
mean transfer: taxa 0.551 · diversity 0.404diversity reversed in 3 of 4 cohorts

Shannon, Simpson, and richness over the 139 model genera (shared reference), logistic model trained on the pooled cohorts. Diversity carries a real signal within populations (training CV 0.639) but does not travel: collapsing to a scalar discards the specific signal while keeping the instability. Within-cohort AUC for n=50 cohorts is noisy; the transfer comparison is the reliable read.

Can a few local labels fix it?

experiment #4 · per-population sign adaptation

Dropping the flippers threw away real signal. The better move is to re-orient them to the target population. You cannot do this without labels, so the honest question is how few: we fit a global model, then let a handful of local labels flip marker directions where the population demands it.

0.320
South Africa read backwards by the global model
0.629
de-inverted with ~40 local labels (ceiling 0.859)
10
marker directions flipped at 40 labels
0.30.50.70.9chance05102040local labels used for calibration0.770.640.630.86
PortugalIndia/MumbaiMexicoSouth Africawithin-cohort ceiling
Deployable form: a calibration artifact

This ships as two parts, not one frozen model: a global model plus a calibrate() step that re-orients it to a new population from a handful of local labels. A reference implementation, not a clinical tool.

artifact data/models/global_logistic.json
code scripts/calibrate.py · demo, calibrate, predict
card docs/Calibration_Model_Card.md

Prior-regularized logistic regression in the 139-genus rank space (a coefficient's sign is a marker's diabetes direction), shrunk toward a global model fit on NHANES+China. Within each cohort, calibration and test samples are disjoint (repeated stratified 5-fold, 12 seeds), so a cohort never scores itself. This interpretable probe is weaker than the flagship CatBoost, so its k=0 transfer values are not the leaderboard AUCs; it isolates the sign mechanism. Local labels are required: there is no unsupervised way to know a population is flipped.

Can it flag its own failure?

experiment #5 · label-free out-of-distribution check

Calibration needs local labels. But could a deployment at least detect, with no labels, that a population is too far from training to trust, and refuse rather than answer backwards? Partly: the most out-of-distribution cohort is exactly the one read backwards.

chance (below = read backwards)0.30.50.78121620Mahalanobis distance from training (label-free)transfer AUCPortugald=8.83India/Mumbaid=8.31Mexicod=10.34South Africad=20.26

Distance is label-free (Mahalanobis of each cohort's centroid from the training distribution in the 139-genus rank space). South Africa, the reversed cohort, sits roughly twice as far out (d=20.26 vs 8.3 to 10.3 for the others; ratios 1.96 to 2.44), so a distance threshold would flag it for refusal. The honest limit: a domain classifier saturates (every cohort is near-perfectly distinguishable from training, AUC 0.986 to 1.000, so it flags everyone including Portugal, which transfers fine), and distance cannot rank the non-reversed cohorts by transfer quality. Detection of a gross reversal is partly possible without labels; fine triage, and correction, still need them.

What eight experiments established

The oral-microbiome diabetes signal is real inside a population, but its direction does not reproduce across studies, so a single pooled model is not merely weaker on a new population, it can be confidently backwards.

And the instability is mostly study-level, not population biology: same-population studies flip nearly as many marker directions as different populations do. The honest target is a global model plus lightweight per-population calibration, not one frozen classifier.

Ruled out (do not re-run)

Sequencing depth as the cause; coarse-graining to diversity; a domain classifier as a triage threshold (it flags everyone); and any fully unsupervised correction (there is none).

Still open

A definitive biology-vs-study call (public data leans study). It needs harmonized, multi-study-per-population data, or a real second African/SCCS cohort, which is data collection and ethics, not analysis.

Does it reproduce?

Live rerun of the preprint's headline scores from the checked-in data. Green = within 0.03 of target.

5-fold CV
0.81
target 0.811 (-0.001)
Portugal holdout
0.698
target 0.698 (+0.000)
3-cohort holdout
0.673
target 0.673 (+0.000)
Reproduces
PASS
within 0.03 on all

Limitations, honestly

What this project is not. Recorded here because a hidden negative is a trap for the next reader.

Small holdout cohorts
Four of six cohorts have 50 to 88 samples, so every AUC carries a wide confidence interval. Trends are more trustworthy than any single number.
No African replication yet
South Africa is the only public African oral cohort with T2D labels, so the reversal cannot yet be pinned as population-biological versus study-specific. This is the most consequential gap.
Associations, not biomarkers
The marker panel is direction-consistent but weak, and several flippers are low-abundance or reagent-associated taxa. These are hypothesis-generating correlations, not validated or causal markers.
Calibration needs local labels
Sign adaptation re-orients a reversed population, but only with labeled local samples. There is no unsupervised fix, and tens of labels only get partway to the within-cohort ceiling.
Genus-level, cross-sectional
16S resolves to genus, collapsing species-level signal, and all data are a single snapshot, so we cannot say whether microbiome shifts precede or follow disease.
Not a clinical tool
NHANES labels are self-reported, and nothing here is validated for diagnosis. Underrepresented populations could receive worse predictions; local calibration and ethics review come before any use.

Roadmap

nowReproducible from public data only

The release must regenerate end-to-end from public sources (NHANES public-use genus tables + SRA cohorts) with no restricted inputs, so every result is independently verifiable.

nowClose the cross-population transfer gap

The core scientific problem: models trained on US+Asian cohorts fall to near-random on unseen populations. Making the oral-microbiome T2DM signal generalize across populations is this project's contribution.

nextBroaden cohort diversity continuously

Generalization needs the populations we lack (Africa, Middle East, Latin America, more Europe). Keep ingesting new public, T2DM-labeled oral-microbiome cohorts as they appear.

nextLiving, citable preprint

Keep docs/Preprint.md current as results evolve, and land a public release with a Zenodo DOI so the work stays discoverable and updatable rather than frozen.

laterCalibrated saliva screening tool

Turn the model into a population-aware, honestly-calibrated T2DM screen (evolve the HTML calculator into a validated tool), the translational end goal.

Studies

active★ flagship
Cross-population oral-microbiome T2DM screening

Transfer varies sharply across unseen cohorts, and in South Africa it reverses: a strong within-cohort signal (CV 0.853) the pooled model reads backwards (transfer 0.320; stratified-bootstrap 95% CI 0.200 to 0.448). Direction, not strength, is what fails to transfer: 71% of informative genera flip diabetes direction across populations. But a within-population decomposition shows this is mostly a study/batch effect, not African biology, four Chinese sub-studies flip 61% of directions among themselves, and the headline marker Atopobium reverses even between two Chinese studies. Restricting to direction-stable genera removes every confident reversal (at an accuracy cost), a few local labels re-orient a flipped population, and a distance gate can flag it label-free. The signal is real but locally oriented, so the honest target is a global model plus per-population calibration, not one frozen classifier.

study #2, TBD

Blog

RSS

Every experiment, honestly: trigger, method, result, interpretation, confidence. Negative results included.

Experiments Log

Running record of experiments on the living project, trigger, method, result, interpretation, confidence. Newest first.


2026-09-25. Audit continuation: uncertainty and the reproduction contract#

Trigger. Resume the September 6 adversarial audit from the Claude session record and test claims the earlier pass did not quantify: per-cohort uncertainty, the geographic narrative, and whether a clean checkout can still regenerate the dashboard after the project-manifest rename.

Evidence. A 10,000-resample stratified bootstrap on the fixed transfer predictions gives:

Cohort Transfer AUC 95% CI Relation to chance
Portugal 0.698 0.544 to 0.837 above
India/Mumbai 0.621 0.455 to 0.782 uncertain
Mexico 0.564 0.441 to 0.684 uncertain
South Africa 0.320 0.200 to 0.448 below

The South African reversal is statistically supported on this test set, while India and Mexico cannot be separated from chance. The ordering does not establish a distance law: the project never defined or tested a geographic, genetic, dietary, or microbiome population- distance variable. One older entry also conflated two models: reversing pooled predictions at AUC 0.320 yields 0.680, not the separately fit within-cohort CV AUC 0.853. Finally, dashboard regeneration failed in a clean workbench because the generator still opened the renamed polaris.toml and its three small extension count tables remained gitignored; its displayed reproduction status was hard-coded to pass. The preprint dataset section also retained a 49-sample/175-total holdout count while the checked-in VSEARCH matrix and results use 50/176, and it described the final training matrix as 10,353 samples (1,378 cases) while the artifact used by the current smoke test contains 10,357 (1,380 cases). The four-sample difference is confined to Shanghai; the older 10,353-sample analyses remain valid as pre-VSEARCH results.

Change. The data contract now computes deterministic transfer intervals and derives its pass/check status from the stated tolerance. The dashboard shows interval whiskers, prints the intervals in tooltips and the cohort table, and calls a cohort above/below chance only when its full interval clears 0.5. Current-facing copy says transfer varies by cohort rather than claiming an unmeasured distance gradient. The generator reads commandem.toml, and the three small public-derived extension count tables are tracked so fresh-clone regeneration is real. The preprint now reports the intervals, corrects the final matrix to 10,357 training samples and the holdouts to 50/176, labels the 10,353-sample LOSO and model-comparison results as pre-VSEARCH, separates the training permutation test from holdout evidence, and distinguishes flipped pooled predictions (0.680) from the within-cohort model (0.853). Older log entries remain intact with inline correction notes.

Confidence. High on the arithmetic, sample counts, generator failures, and conditional test-set intervals. The intervals condition on this one study sample and do not capture study- selection, label, preprocessing, or model-development uncertainty; the absence of a tested distance relationship is a scope correction, not evidence that distance has no effect.


2026-09-06. Adversarial audit: reproduction confirmed, two mechanism claims downgraded#

Trigger. An independent adversarial audit of the full research chain (experiments #1-#8, preprint, data contract), run before any release decision. Older entries are NOT rewritten; they carry short correction notes pointing here, so the record stays honest about what we believed when.

What the audit VERIFIED. All fast experiments reproduce exactly (re-run produced zero diff against the committed JSONs; baseline CV 0.81 / Portugal 0.698 / 3-cohort 0.673). The 71% flip prevalence recomputes by hand (57 of 80). The LOCO de-inversion, the sign-adaptation curve and its 0/4/5/7/10 flip counts, the OOD numbers, diversity-transfers-worse, and the 6x prevalence trend all check out. The 0.853 / 0.859 / 0.79 within-SA values are consistent (different models, each labeled correctly).

Finding 1 (downgrade). The "top-20 markers flip" table does not test the model's top markers. experiment_marker_stability.py selects the top 20 by |mean point-biserial r| in the pooled training populations. By that metric Atopobium ranks 133 of 139 (NHANES +0.02, China -0.04), so the marker the narrative is about is not in the table. CatBoost importance (where Atopobium is rank 1) and linear association strength are different rankings; the entries conflated them. The table is also non-monotone: India and South Africa share 40% flips at transfer 0.621 vs 0.320. The flip-fraction-tracks-transfer claim is SUGGESTIVE for association-top markers, and is no longer presented as the Atopobium mechanism.

Finding 2 (downgrade). Atopobium's "training direction -0.401" is a pooling artifact. The -0.401 was computed on the combined NHANES+China pool, whose composition confounds it: the Chinese studies are ~82% T2D (clinics) and NHANES ~9% (survey), so a genus that differs by study/specimen inherits a spurious pooled association with the label (Simpson's paradox). Within each study Atopobium's association is near zero (NHANES +0.02, China -0.04), and experiment #8 later showed it flips between two Chinese studies (+0.19 vs -0.46). There is no settled biological "training direction" for Atopobium; the within-SA +0.47 stands, but the -0.40 to +0.47 "flip" compared an artifact to a signal. This strengthens, not weakens, the experiment #8 conclusion that the instability is study-level.

Finding 3 (disclosure). The 61% vs 71% comparison has nested structure. The four Chinese sub-studies appear both as the within-China groups and, pooled, inside the cross-population set, so the two estimates are not independent, and their denominators differ (99 vs 79 genera informative in >=2 groups). The direction of the conclusion is unchanged; the comparison is descriptive, not a clean decomposition.

Finding 4 (precision). "~2.4x farther" for South Africa's Mahalanobis distance was the most favorable pairwise ratio. The range against the other cohorts is 1.96x (vs Mexico) to 2.44x (vs India), ~2.2x on average. Also, the domain classifier range is 0.986 to 1.000, so "0.99+" was imprecise for Portugal (0.986). And in sign_adaptation.json Portugal carries reversed=true because the LOGISTIC probe's k=0 transfer is 0.483; that flag describes the probe, not the flagship CatBoost (0.698), and should not be read as "Portugal reverses."

Net effect on the story. The core chain holds: the SA reversal is real (not depth), 71% of informative genera flip, restriction de-inverts safely, a few labels re-orient, detection is partly label-free, and the instability is mostly study-level. What is retired is the specific "Atopobium's training direction flips" mechanism; the honest mechanism evidence is the broad association anti-correlation (r = -0.237) plus Atopobium's high model importance and its within-China instability. Preprint sections 6, 6.1, 6.4, and 6.7 were revised accordingly.


2026-09-02. Experiment #8: population biology or study effect? Mostly study#

Trigger. Every prior experiment brackets the South Africa reversal but none says whether it is African population biology or a study-specific artifact. A second African cohort would settle it but none is public. Desk-feasible proxy: the training pool holds four Chinese sub-studies (anhui, shandong, sichuan, shanghai), same population, different labs and runs. If directions flip among same-population studies about as much as across populations, then study/batch alone reverses signs and the population-biology reading weakens.

Method. Per-genus point-biserial T2D association within each group. Two comparisons, plus an Atopobium spotlight: association-vector correlation (as in the 2026-08-27 train-vs-SA analysis), mean over within-China study pairs vs cross-population pairs; and sign-flip rate among genera informative (|r| >= 0.15) in >= 2 groups. scripts/experiment_population_vs_study.py.

Result.

  • Sign-flip rate: 61% within China (of 99 genera, same population, 4 studies) vs 71% across populations (of 79). Same-population studies flip nearly as often as entirely different ones. [Correction 2026-09-06: the comparison is nested, the four Chinese sub-studies also appear, pooled, inside the cross-population set, and the denominators differ (99 vs 79 genera), so the two rates are descriptive rather than an independent decomposition. Direction unchanged. See the audit entry.]
  • Association-vector correlation: +0.163 within-China vs -0.001 cross-population. Population is not nothing (crossing populations collapses the weak agreement to zero), but same-population concordance is already weak.
  • Atopobium flips within China too: anhui +0.19, shandong -0.46, sichuan -0.25, shanghai -0.14, a 0.65 swing inside one population, comparable to its cross-population range (SA +0.47).

Interpretation. The direction instability is mostly a study/batch phenomenon, not primarily population biology. Same-population studies already disagree on ~61% of directions; crossing populations adds only ~10 points. So the South Africa reversal is not distinguishable from generic study-level instability, and the tempting "African biology" explanation is not supported by public data. A weak population signal does survive (within-China correlation +0.16 exceeds cross-population 0.00), so population contributes a real but secondary increment.

Confidence. The within-vs-cross gap is consistent across both metrics and the Atopobium within-China flip is unambiguous. GRADED: the Chinese sub-studies are small (n=45 to 347) so their per-genus signs are themselves noisy; some within-China flipping is small-sample noise rather than batch per se, but that unreliability is exactly the finding, directional associations do not reproduce at these sample sizes, within or across populations.

Implication. Reframes the flagship. The transfer gap is dominated by non-reproducible, study-level direction estimates rather than clean population biology, which raises the bar for any biomarker claim and makes harmonized, multi-study-per-population data (not just more populations) the real prerequisite. It also means the SA reversal should be presented as a study-level effect first, with population as a weaker contributor.


2026-09-01. Experiment #7: going coarse does not help, community diversity transfers worse than taxa#

Trigger. Experiments #2-#6 show the taxonomic signal is real but direction-unstable across populations. Maybe a coarser summary, alpha diversity, is weaker but more portable, because a single scalar is less sensitive to which specific genera flip.

Method. Over the 139 model genera (a shared reference across cohorts) compute per-sample Shannon, Simpson, and richness from raw counts. A logistic model on the three diversity features is trained on the pooled cohorts, cross-validated within each cohort, and transferred to each unseen cohort, then compared to the 139-genus taxonomic model. Training pool and Portugal come from the unified raw-count table joined to the pkl sample-id labels; India/Mexico/SA from their interim counts. scripts/experiment_diversity_transfer.py.

Result.

Cohort diversity within-CV diversity transfer taxa transfer Shannon (T2D - ctrl)
Portugal 0.27 0.432 0.698 +0.068
India/Mumbai 0.50 0.544 0.621 -0.093
Mexico 0.62 0.451 0.564 +0.156
South Africa 0.79 0.188 0.320 +0.056
  • Diversity transfers WORSE than taxa (mean 0.404 vs 0.551) and reverses in 3 of 4 cohorts (taxa reverse in 1). South Africa reverses even harder under diversity (0.188 vs 0.320).
  • The direction of diversity itself is population-specific: Shannon is higher in T2D in most cohorts but lower in India, so the sign flips like the taxa do.
  • Diversity does carry a within-population signal (training CV 0.639; South Africa within 0.79), it just does not travel.

Interpretation. Coarse-graining is not a shortcut to portability. The direction instability is baked into the community-disease relationship at every level, right down to a single number: collapsing to diversity throws away the specific taxonomic signal while keeping the instability. This refutes the "coarse features are more portable" hypothesis.

Confidence. Clean comparison on shared features. GRADED DOWN: within-CV AUC for n=50 cohorts is noisy (Portugal 0.27 is a small-sample artifact, not a real anti-signal); the reliable read is the transfer comparison, where diversity is clearly worse and reverses more.

Implication. No free lunch from summary statistics. Portability has to come from modeling direction (experiments #3, #4), not from a coarser feature space.


2026-09-01. Experiment #6: what makes a marker portable? Stable markers are the prevalent ones#

Trigger. Experiment #2 split the 139 genera into direction-stable markers and sign-reversing flippers but did not ask why. If flippers are just the rare, low-abundance taxa whose association sign is noisy, a portable panel should favor prevalent, abundant genera, and the prevalent-yet- reversing genera are the interesting exceptions.

Method. For each genus compute two label-free properties from the full raw count table: prevalence (fraction of samples detected) and mean relative abundance. Join with the stability class from experiment #2; compare stable markers vs flippers (Mann-Whitney), fit a small logistic (flipper ~ log-prevalence + log-abundance), and flag the prevalent+abundant flippers. scripts/experiment_marker_properties.py.

Result.

  • Direction-stable markers are about 6x more prevalent than flippers (median 21% of samples vs 3.6%); abundance shows the same direction. The flipper-vs-stable logistic loads most on low prevalence (standardized coef -0.34).
  • But the difference is not statistically significant (prevalence p=0.13, abundance p=0.12), with only 15 stable genera, so it is a suggestive trend, not a proven rule.
  • A distinct set of prevalent, abundant genera reverse anyway: Faecalibacterium, Veillonella, Akkermansia, Actinomyces, Alloprevotella, Atopobium, Campylobacter, and more. These reverse for reasons other than sparsity and are the biologically meaningful reversers.

Interpretation. Rarity explains part of the flipping (rare taxa have noisier association signs), so a portable panel should prefer prevalent, abundant genera. But prevalence is no guarantee: the dominant markers a model relies on (Atopobium) can be both common and reversing, which is exactly why the pooled model fails confidently rather than quietly.

Confidence. The prevalence gap is a clear 6x effect but underpowered (p0.12-0.13); the prevalent-flipper list is descriptive. Properties are computed on the training-pool raw counts.

Implication. Actionable for panel design: start from prevalent, abundant genera, but test each marker's direction per population rather than trusting prevalence alone.


2026-09-01. Experiment #5: can the model flag its own failure, without labels?#

Trigger. Experiment #4 re-orients a reversed population but needs local labels. The open question is detection rather than correction: before any labels, can a deployment tell it is out-of-distribution (OOD) and likely to answer backwards, and refuse instead of being confidently wrong?

Method. Reference = the pooled training distribution in the 139-genus rank space. Two label-free OOD scores per unseen cohort: (1) a domain-classifier AUC (5-fold CV AUC separating training from the cohort, balanced downsample, 15 seeds), and (2) the Mahalanobis distance of the cohort centroid from training (Ledoit-Wolf shrinkage covariance). Then ask whether either score, computed with no diabetes labels, flags the reversed cohort. scripts/experiment_ood_detection.py.

Result.

Cohort Transfer Domain-classifier AUC Mahalanobis
South Africa (reversed) 0.320 1.000 20.3
Mexico 0.564 0.998 10.3
India/Mumbai 0.621 0.996 8.3
Portugal 0.698 0.986 8.8
  • The domain classifier saturates: every cohort is ~perfectly distinguishable from training (0.986 to 1.000), because batch and study effects make any new cohort trivially OOD. Its ranking is perfectly monotonic with transfer (Spearman -1.0 over the four cohorts) but the spread is only 0.014, so a threshold would refuse everyone, including Portugal (transfers fine).
  • Mahalanobis distance separates cleanly: South Africa's centroid sits ~2.4x farther out (20.3) than any other cohort (8 to 10). The reversed cohort is the distributional outlier. [Correction 2026-09-06: 2.4x was the most favorable pairwise ratio; the range is 1.96x (vs Mexico) to 2.44x (vs India), ~2.2x on average. See the audit entry.]

Interpretation. Detection of a gross reversal is partly possible without labels: a distance threshold flags South Africa for refusal, so a deployment need not answer backwards. But two honest limits remain. A domain classifier alone cannot triage (batch effects make everything look OOD), and even the usable distance signal cannot rank the non-reversed cohorts by transfer quality or distinguish attenuation from reversal among them. Being different is only loosely tied to being wrong.

Confidence. Four unseen cohorts, so any OOD-versus-transfer correlation is weak and the Spearman -1.0 rests on tiny domain-AUC gaps; the robust part is the Mahalanobis outlier (South Africa 2.4x the rest), which is a clear, single-number effect. The claim is deliberately modest: partial unsupervised detection, not reliable triage.

Implication. A safe deployment can pair a label-free distance gate (refuse or down-weight grossly OOD populations) with the label-based calibration of experiment #4 (re-orient once a few local labels arrive). Detection buys caution; correction still buys accuracy.


2026-08-28. Experiment #4: per-population sign adaptation, a few local labels re-orient the reversed model#

Trigger. Experiment #3 de-inverted South Africa by dropping the sign-reversing genera, but that throws away real (if flipped) signal and costs accuracy. The better move is to ADAPT the marker directions to the target population rather than discard them. Pure unsupervised sign correction is impossible (you cannot know a population is flipped without labels), so the honest question is few-shot: with k labeled local samples, can we re-orient the global model and recover the within-cohort signal it reads backwards?

Method. scripts/experiment_sign_adaptation.py. Interpretable logistic regression in the 139-genus rank space, where a coefficient's sign is a marker's diabetes direction. A global model is fit on the NHANES+China pool (beta_global). For a target cohort we fit a prior-regularized logistic on k local samples: minimize logloss + (lambda/2)||beta - beta_global||^2, so at k=0 it is the global model and as k grows coefficients can flip sign where the population demands it, shrunk toward the global prior for stability. Leakage guard: within each cohort, calibration and test samples are disjoint (repeated stratified 5-fold, 12 seeds); the global model never saw the cohort. We sweep k in {0, 5, 10, 20, 40} and count coefficient sign flips vs the global model.

Result (AUC by number of local labels):

Cohort k=0 k=10 k=20 k=40 within-cohort ceiling
South Africa 0.320 0.405 0.487 0.629 0.859
Portugal 0.483 0.511 0.517 0.562 0.766
India/Mumbai 0.545 0.557 0.573 0.587 0.639
Mexico 0.620 0.619 0.626 0.633 0.627
  • South Africa, read backwards at k=0 (0.320), climbs monotonically and crosses chance by ~40 local labels (0.629), heading toward its within-cohort ceiling of 0.859.
  • The mechanism is visible: South Africa's coefficient sign flips vs the global model grow 0, 4, 5, 7, 10 across the k grid. Adaptation is literally re-orientation of markers.
  • Cohorts that were not reversed gain little (Mexico has no signal to recover; India and Portugal inch up), which is the expected pattern.

Interpretation. The reversal is fixable, but only with local supervision. A small calibration set re-orients the flipped markers and turns a confidently-wrong model into a useful one, which is exactly "per-population sign adaptation." It also sets a rough cost: on the order of tens of labeled local samples to de-invert South Africa, not hundreds.

Confidence. The de-inversion trend and the growing sign-flip count are robust (leakage-free splits, 12 seeds). GRADED DOWN: this is an interpretable logistic probe, weaker than the flagship CatBoost, so its k=0 transfer values are not the leaderboard AUCs (Portugal's logistic transfer is 0.483 here versus 0.698 for CatBoost); the point is the mechanism and the trend, not the absolute level. Even at k=40, South Africa (0.629) is well short of its 0.859 ceiling, so tens of labels help but do not fully recover the signal, and per-cohort n is small.

Implication. The deployable design is a global model plus a lightweight local calibration step, not a single frozen classifier. Any future African/SCCS cohort gives a direct test: does the same handful of local labels re-orient it the same way?


2026-08-28. Experiment #3: portability over accuracy, a stable-feature model stops failing backwards#

Trigger. Experiment #2 showed the transfer gap is a direction problem: 71% of informative genera flip sign across populations and the model's top markers flip 40 to 50% in distant cohorts. The convicted hypothesis: a model restricted to the direction-stable genera should transfer more honestly (fewer confident reversals), even at the cost of raw accuracy, because it drops the flippers the pooled model reads backwards.

Method. scripts/experiment_portability.py. Two tests, both leakage-free.

  • A. Stable-restricted vs full transfer. The stable set from experiment #2 was defined using the test cohorts, so restricting features and testing on those same cohorts would be circular. So we select leave-one-cohort-out: to evaluate transfer to cohort T, the direction-stable genera are chosen from the OTHER five cohorts only, so T's labels never inform its own feature set. Training always uses only the NHANES+China pool. Compare the stable-restricted CatBoost to the full 139-feature model on each held-out cohort.
  • B. Mechanism ablation. Drop Atopobium, the top marker chosen by TRAINING importance alone (leakage-free), and measure South Africa transfer.

Result.

  • Restriction trades accuracy for portability. Portugal falls 0.698 to 0.589 (the full model is better where transfer already works), India holds (0.621 to 0.616), Mexico edges up (0.564 to 0.600), and South Africa de-inverts: 0.320 to 0.515, no longer read backwards.
  • Reversed cohorts: 1 to 0. Mean transfer rises modestly, 0.551 to 0.580.
  • Ablation: dropping Atopobium alone moves South Africa 0.320 to 0.410, a partial fix, so the reversal is broad and not carried by a single marker (consistent with the 2026-08-27 finding that 7 of the top 12 genera flip).

Interpretation. You cannot build a more accurate global model this way, but you can build a safer one: restricting to direction-stable genera removes the confident reversal and makes the model fail to chance rather than fail backwards. This is the concrete payoff of "portability over accuracy," and it validates experiment #2's thesis with a leakage-free causal test: the flippers are what break transfer, and removing them fixes the break at a measurable accuracy cost.

Confidence. The de-inversion and the reversed-cohorts 1-to-0 result are robust to leakage (LOCO selection, training on the training pool only). GRADED DOWN: a move toward 0.5 is a loss of confident-wrongness, not a gain of skill; per-cohort n is small (50 to 88); and the mean-transfer gain (0.03) is modest and within noise for cohorts this size. The honest headline is safety, not accuracy.

Implication. The screening target should be a portability-first model (direction-stable features, or per-population sign adaptation), not a single global classifier optimized for AUC. Any future African/SCCS cohort tests whether the same stable set keeps South Africa non-reversed.


2026-08-28. Experiment #2: which genera generalize? Direction, not strength, fails to transfer#

Trigger. The South Africa reversal showed Atopobium flips its diabetes direction (control-associated in the training populations, case-associated in SA). That convicts a systematic question across ALL six populations: which oral genera keep a consistent T2D direction (trustworthy markers) and which flip? A model can only transfer on the stable ones.

Method. Six populations placed in the shared 139-genus space: NHANES (US) and China from the training pool (split by study_labels), Portugal from the geographic holdout, and India/Mumbai, Mexico, South Africa as external cohorts, each per-sample rank-transformed exactly as at inference (leakage-free). Per genus per cohort: point-biserial correlation between the sample's rank for that genus and the T2D label (sign = direction, |r| = strength). Stability is defined by DIRECTION consistency, not magnitude: NHANES (n=9683) has a diffuse signal (median |r| = 0.01), so a fixed magnitude gate would wrongly mark the largest cohort uninformative. A cohort casts a sign vote once |r| >= 0.05; a genus is "robust" when its sign agrees across >= 5 of 6 cohorts, a "flipper" when it clears |r| >= 0.15 in both directions. scripts/experiment_marker_stability.py.

Result.

  • 71% of genera informative (|r| >= 0.15) in >= 2 populations reverse sign across them (57 of 80). This figure excludes NHANES, so it does not depend on the magnitude artifact.
  • Only 15 genera hold a single direction across >= 5 populations, and all are weak (|mean r| 0.03 to 0.12). Strongest stable case-associated: Streptococcus (+0.12), Peptoniphilus, Enterococcus, Bifidobacterium. Stable control-associated: Fretibacterium, Johnsonella, Treponema, Cloacibacterium.
  • 56 flippers, Atopobium among them (NHANES +0.02, China -0.04, Portugal -0.07, India +0.15, Mexico -0.19, SA +0.47). Others: Faecalibacterium, Campylobacter, Bacteroides, Actinomyces.
  • Effect: marker direction predicts transfer. Of the pooled model's 20 strongest training markers, the fraction that flip sign tracks how well the model transfers: Portugal 10% (transfer AUC 0.698), India 40% (0.621), Mexico 50% (0.564), South Africa 40% with the highest-importance marker among them (reversed to 0.320). [Correction 2026-09-06: "strongest training markers" here means strongest by pooled ASSOCIATION, not by model importance; Atopobium ranks 133/139 by that metric and is not in this set, and the relation is non-monotone (India = SA = 40%). See the audit entry.]

Interpretation. The cross-population transfer gap is a direction problem, not a strength problem. The diabetes signal is real and often strong inside each population, but the SIGN of almost every genus-to-diabetes association is population-specific. South Africa is not an outlier so much as the extreme of a pervasive instability. This is why a pooled model degrades smoothly with population distance and then breaks entirely on SA: it is averaging directions that do not agree. The few direction-stable genera are too weak to carry a portable model on their own.

[Correction 2026-09-25: no population-distance variable was defined or tested. The supported claim is that transfer varies sharply by cohort, with a statistically supported reversal in South Africa; a smooth distance relationship is retired. See the audit entry.]

Confidence. The 71% reversal prevalence and the marker-flip-vs-transfer effect are robust to the NHANES magnitude artifact (the effect uses only sign, the prevalence excludes NHANES). GRADED DOWN: these are associations in modest per-cohort samples (n = 50 to 88 for four of six cohorts), not causal or diagnostic claims, and several flippers are low-abundance or environmental/reagent-associated taxa (Serratia, Burkholderia, Delftia, Paenibacillus) whose signs may reflect contamination rather than biology. The "robust" panel is direction-consistent but weak, not a validated biomarker set.

Implication. Points the modeling work at portability rather than accuracy: a model built on direction-stable genera, or one that adapts sign per population, is the honest path to transfer, and any future African/SCCS cohort gives a direct test of whether the flippers (Atopobium first) flip the same way. It also reframes the flagship claim: the oral-microbiome T2D signal is real but locally oriented, so a single global screening model is the wrong target.


2026-08-27. The SA reversal is broad and led by Atopobium (the model's top marker)#

Trigger. South Africa reverses (strong within-cohort signal, backwards transfer). Which genera flip, and is it a few markers or a global inversion?

Method. For each of the 139 model genera, point-biserial correlation of its (rank-transformed) abundance with the T2D label, computed in the TRAINING populations and in South Africa; then correlate the two per-feature association vectors and inspect the model's most important features. scripts/experiment_feature_reversal.py.

Result.

  • Per-feature T2D associations, training vs SA: r = -0.237 (p = 0.005), a significant broad anti-correlation.
  • The model's dominant marker Atopobium (importance ~19, several times the next) flips hard: assoc_train -0.401 (lower in diabetics) to assoc_SA +0.467 (higher in diabetics). [Correction 2026-09-06: the -0.401 is a POOLED NHANES+China estimate confounded by study composition (China ~82% T2D vs NHANES ~9%); within each study Atopobium's association is near zero (+0.02 / -0.04). The within-SA +0.47 stands; the "training direction" does not. See the audit entry and experiment #8.]
  • 7 of the top 12 model genera flip diabetes-direction in SA (Atopobium, Phocaeicola, Alloprevotella, Treponema, Neisseria, Mogibacterium, Rhodobacter).

Interpretation. The reversal is broad (significant negative association-correlation) and concentrated in the genera the model relies on most, led by Atopobium. Because Atopobium alone carries ~19% importance and its diabetes association strongly flips, the pooled model is confidently backwards on SA. It is not a single-feature artifact: most of the model's reasoning inverts. Notably, Atopobium is the preprint's headline T2D marker, so its reversal in South Africa is exactly the finding to scrutinize.

Confidence. The Atopobium flip and the significant negative correlation are clear. Whether this reflects population biology or a SA-study batch effect is still unresolved, a systematic case/control batch difference could also flip associations broadly.

Implication. Names the mechanism (Atopobium-led broad reversal) and gives a concrete, pre-registered thing to check in any future African or SCCS cohort: does Atopobium's diabetes direction flip there too? If it does, population biology gains support; if not, the SA study is idiosyncratic.


2026-08-27. No second public African oral cohort to replicate the reversal#

Trigger. South Africa shows a strong, reversed T2D signal. To decide whether that is population biology or specific to one study, we need an independent African (or African-descent) oral 16S cohort with diabetes labels.

Method. Systematic search across NCBI SRA/BioProject, ENA, GEO, Qiita, PubMed, Google Scholar, and African-microbiome/H3Africa angles (2019-2026), filtering for oral (not gut) 16S with per-subject T2D or glycemic labels.

Result. No clean match exists. Public African oral 16S with T2D labels is essentially a single cohort, the one we already have (PRJNA723337). Nearest options, each caveated:

  • Uganda PRJNA1087153 (public): gingival crevicular fluid, Nanopore full-length 16S, n=45 (26 diabetes / 19 non-diabetes), diabetes diagnosis but no HbA1c. Different specimen and platform, so only a weak directional check, not a like-for-like replication.
  • SCCS (African-American, US): mouth rinse, Illumina V4, n=294, real T2D labels, the methodologically closest resource, but CONTROLLED ACCESS (requires a data-use request).
  • Ruled out: Nigeria (samples pooled, no deposition), Egypt PRJNA1475433 (healthy-only), Angola/Zimbabwe saliva (no diabetes labels). A "SCCS on SRA" lead was a mislabel, the accession is a rhesus-macaque project.

Interpretation. The population-vs-study question CANNOT be settled with off-the-shelf public data today. This is a genuine data-availability limitation, not a failure to search: African oral-microbiome + diabetes data is thin (most African diabetes-microbiome work is gut/stool).

Confidence. High on the availability conclusion (convergent search across repositories).

Implication. Paths forward: (a) run the Uganda cohort as an explicitly weak, cross-platform directional check; (b) apply for SCCS controlled access, the rigorous answer; (c) partner for new data (e.g. H3Africa-linked oral cohort). Until then, the SA reversal is reported as a real, verified, single-cohort finding whose population-generality is unresolved.


2026-08-27. South Africa's inversion is a REAL, reversed signal (not noise)#

Trigger. Depth was ruled out. Is SA's inversion (AUC 0.320) a genuine reversed signal, or just small-sample noise that happens to land below 0.5?

Method. For each new cohort, in the SAME 139-feature rank space, compare the TRANSFER AUC (pooled model scored on the cohort) with a WITHIN-COHORT AUC (5-fold CV of a model trained and tested inside the cohort). scripts/experiment_within_cohort.py.

Result.

Cohort n Transfer AUC Within-cohort CV
India/Mumbai 60 0.621 0.638
Mexico 87 0.564 0.566
South Africa 88 0.320 0.853

Interpretation. Three distinct regimes:

  • South Africa has a strong, learnable T2D signal (within-CV 0.853) that is reversed relative to the training populations, which is exactly why the pooled model lands at 0.320. Flip the predictions in SA and the same model would score ~0.85. [Correction 2026-09-25: flipping these pooled-model scores gives AUC 0.680 (= 1 - 0.320), not 0.85. The 0.853 figure is a separately trained within-cohort CV model; it shows learnable local structure but is not the flipped pooled model. See the audit entry.] So the inversion is real signal pointing the opposite way, not noise and not depth.
  • Mexico is ~chance both ways (0.566 within, 0.564 transfer): little learnable T2D signal at all, consistent with its 100%-periodontitis composition masking the contrast.
  • India transfers about as well as it is internally learnable (0.638 vs 0.621).

So "cross-population failure" is not one thing: it has at least two modes, no signal (Mexico) and reversed signal (South Africa).

Confidence. The SA reversal is clear: within-CV 0.853 is far above chance, so even allowing for small-n CV optimism there is strong, real signal that the pooled model gets backwards. What one African cohort CANNOT separate: population-biological reversal vs a study/batch-specific reversal (SA is a single study).

Implication. The oral-microbiome-to-T2D relationship can genuinely differ, even reverse, across populations. A pooled model is not merely unreliable across populations, it can be confidently wrong, which argues for population-specific calibration (or sign checks) before any deployment. Next: a second African (or African-descent) cohort to test whether the reversal reproduces (population) or is specific to this study.


2026-08-27. South Africa inversion is NOT a sequencing-depth artifact#

Trigger. SA transferred at AUC 0.320 (inverted). The leading suspect was sequencing depth: SA is deep (~218k mapped reads/sample) vs the shallower cohorts, and deep sequencing detects more rare genera → a denser per-sample rank profile that could be out-of-distribution for the training data.

Method. Rarefied each SA sample's genus-count vector to a ladder of shallower depths (50k → 1k, multinomial subsample, 3 seeds averaged), re-applied the identical per-sample rank transform, and re-scored with the pooled-cohort CatBoost. scripts/experiment_sa_rarefaction.py.

Result.

SA depth Transfer AUC
full (~218k) 0.320
50,000 0.358
20,000 0.406
10,000 0.376
5,000 0.380
2,000 0.369
1,000 0.377

Total AUC swing across depth: 0.048, and it stays inverted (well below 0.5) at every depth, down to 1,000 reads.

Interpretation (hypothesis rejected). Sequencing depth is not the driver of the inversion. Rarefying to matched-or-shallower depth nudges AUC up by at most ~0.05 but never uncrosses 0.5. So the model's T2D-associated oral-microbiome features genuinely anti-correlate with T2D in this African cohort, a real cross-population reversal, not a technical/depth artifact.

Confidence. Solid rejection of the depth hypothesis (consistent across six depths × three seeds). What remains open is why it inverts, population biology vs a study/batch confound vs cohort composition, which one deep cohort can't separate.

Implication. Strengthens the core thesis: cross-population failure is real and can be directional, not just noisy. Next suspects to test: study/batch effect (SA is a single-study cohort) and whether a second African/saliva cohort reproduces the sign.


2026-08-23. South Africa (PRJNA723337): signal INVERSION (AUC 0.320, verified)#

Trigger. Third plaque cohort, completing the geographic sweep (Africa gap, the only public African oral 16S T2DM data).

Method. PRJNA723337, 88 subgingival-plaque samples (57 diabetic [DM+Known_DM] / 31 control [Normal]; Pre-DM excluded). Single-end, deep (~227k reads/sample). Required --fastq_qmax 50 (SA reads hit Q43, above vsearch's default 41, a real pipeline fix, now in the generic processor). Labels verified against NCBI BioSample isolate (spot-checked: DM_13→diabetic, Known_DM_23→diabetic, etc., all correct). Same 139-feature → rank → CatBoost(existing pool) transfer test.

Result (verified inversion).

value
Transfer AUC 0.320 (below random)
Flipped-label AUC 0.680 (systematic, not noise)
Mean pred: diabetic vs control 0.829 vs 0.884 (controls scored more diabetic)
Genera observed 125/139 (highest, deep sequencing)

Four-cohort gradient (all new cohorts this session):

Population Specimen Depth Transfer AUC
European (Portugal) saliva n/a 0.698
S. Asian (India) plaque ~97k 0.621
Latin Am. (Mexico) saliva ~50k 0.564
African (South Africa) plaque ~227k 0.320 (inverted)

Interpretation. A real, verified systematic inversion, the extreme end of cross-population failure: the model's T2DM-associated features anti-correlate with T2DM in South Africa. The gradient above roughly tracks population distance from the US+Asian training set. BUT the cause is entangled and must not be overclaimed as biological: SA differs on four axes at once, population (African), specimen (plaque), sequencing depth (227k reads → far denser genus detection → a very different rank-profile distribution than the shallow training cohorts), and clinic composition. The compressed, uniformly-high predictions (0.83–0.88) point to a severe distribution shift where the rank transform interacts badly with the depth difference.

[Correction 2026-09-25: the apparent ordering is descriptive; no population-distance measure or distance-transfer test was specified. Current-facing surfaces no longer call it a gradient.]

Confidence. Inversion is verified and systematic (labels correct, flip=0.68). Attribution is NOT established, do not claim "African oral T2DM signal is inverted."

Implication. (1) Naive cross-population application isn't just unreliable, it can be actively wrong, a strong argument for mandatory population-specific calibration before any deployment. (2) New concrete suspect: sequencing-depth harmonization. Testable next step, rarefy/normalize SA to comparable depth before the rank transform and re-run; if the inversion softens, depth (not population) is a major driver.


2026-08-23. Mexico (PRJNA1208116) transfer test REFUTES the specimen hypothesis#

Trigger. India (plaque) transferred only partially and naive pooling hurt; the leading hypothesis was that specimen mismatch (plaque vs saliva) was the culprit. Mexico is saliva, a specimen-matched, balanced, new-continent test to isolate population effect from the plaque confound.

Method. PRJNA1208116, 87 saliva samples (41 T2DM / 46 control; labels recovered from the paper's Supplementary Table S1, CC-BY). Same VSEARCH → 139-feature → per-sample rank → CatBoost(existing pool) transfer test. scripts/integrate_cohort.py.

Result.

Holdout Specimen Population AUC
Portugal saliva European 0.698
3-cohort mixed mixed 0.673
India plaque S. Asian 0.621
Mexico saliva Latin American 0.564
Asian↔US (known) n/a n/a ~0.541

Mexico had the best feature overlap (109/139) yet transferred worst, near random.

Interpretation (hypothesis refuted). Specimen match did NOT rescue transfer: saliva Mexico (0.564) transferred worse than plaque India (0.621). So specimen type is not the dominant barrier. BUT the result is confounded: Mexico is a 100%-periodontitis cohort (Severe 48 / Moderate 31 / Mild 21), so the T2DM contrast is within periodontitis, where the diabetes signal is subtle relative to the dominant periodontitis-driven microbiome shift, and this differs from the general-population (NHANES-heavy) training set. Two entangled explanations remain: (a) population distance (Latin American), (b) cohort composition (all-periodontitis clinic case-control vs general-population survey). One cohort cannot separate them.

Confidence. Result solid (near-random, n=87, balanced). Interpretation graded: specimen-hypothesis refutation is well-supported; population-vs-periodontitis remains open.

Implication (reframes the roadmap). The clinic-recruited, periodontitis-focused cohorts (India, Mexico) test a different, harder, confounded task than general-population T2DM screening. To measure pure population transfer, prioritize general-population survey cohorts (like NHANES itself, or the Qatar Biobank) over periodontitis clinic case-controls. This is now the sharper dataset-search criterion.


2026-08-23. Does pooling India into training help the existing holdouts? (NO)#

Trigger. After the India transfer test, ask the complementary question: does adding India to the training pool improve generalization to the current holdouts?

Method. Baseline CatBoost (existing pooled cohorts) vs. the same with India's 60 plaque samples appended to training (rank-transformed into the 139-feature space). Re-scored the Portugal and 3-cohort holdouts. scripts/experiment_india_in_pool.py.

Result (negative).

Model Portugal 3-cohort
baseline (no India) 0.698 0.673
+ India pooled 0.624 (−0.074) 0.621 (−0.052)

Interpretation. Naive pooling of the India plaque cohort degrades transfer to the (saliva) holdouts by 0.05–0.07 AUC. The specimen mismatch (plaque vs saliva) plus partial feature overlap acts as a domain shift that adds noise rather than useful diversity. Consistent with the project's recurring finding that naive cross-study pooling hurts generalization (prior domain-adaptation attempts also degraded holdout).

Confidence. Clear directional result, both holdouts drop materially and in the same direction. Caveat: single cohort, n=60, specimen-confounded.

Implication. Don't naively pool across specimen types. India (plaque) needs site-matching or explicit domain handling; specimen type is a confound to control before added geography can help. Motivates seeking a saliva cohort for new geography.


2026-08-23. India/Mumbai (PRJNA1240053) cross-population transfer test#

Trigger. Living-project goal "close the cross-population transfer gap": test whether the existing pooled model generalizes to a never-seen population.

Method. Downloaded the India/Mumbai plaque 16S cohort (60 samples, labels from SRA aliases: 40 diabetic [T2DM + T2DM+perio] / 20 non-diabetic [healthy + perio]). Processed through the same VSEARCH pipeline (merge → maxee 1.0/minlen 200 filter → 97% closed-ref vs HOMD v16.03 → genus). Mapped to the 139 model genera, applied the identical leakage-free per-sample rank transform, and scored with CatBoost (depth 4, iters 200, lr 0.1, balanced, seed 42) trained on the existing pooled cohorts. India was never seen in training, a clean LOSO-style holdout. Scripts: scripts/process_india_vsearch.py, scripts/integrate_india.py.

Result.

Holdout AUC n
India/Mumbai (new) 0.621 60
Portugal (ref) 0.698 50
3-cohort combined (ref) 0.673 176
Asian↔US (known failure) ~0.541 n/a

Only 76/139 model genera were observed in India (plaque community + region differences); the 63 absent genera tie at the bottom rank and add little signal.

Interpretation. Partial transfer, above the near-random cross-continental failure floor (0.541) and above chance, but below the European holdouts. The model retains some signal on a new population, which is notable given three stacked disadvantages: (1) never-seen population, (2) plaque, not saliva (specimen mismatch), (3) only 55% feature overlap. The plaque mismatch likely understates population transfer, so 0.621 is a conservative floor for "does the signal reach India."

Confidence. Suggestive, not definitive, n=60, single cohort, specimen mismatch. Not strong enough to claim India generalization; strong enough to justify adding saliva Indian data and testing whether site-matching lifts it.

Next. (a) Get a saliva Indian cohort to separate site effect from population effect. (b) Add India to the training pool and re-check whether it helps other holdouts (diversity gain). (c) Pending Mexico/South Africa labels for more geography.