Experiments Log
Running record of experiments on the living project, trigger, method, result, interpretation, confidence. Newest first.
2026-09-25. Audit continuation: uncertainty and the reproduction contract#
Trigger. Resume the September 6 adversarial audit from the Claude session record and test claims the earlier pass did not quantify: per-cohort uncertainty, the geographic narrative, and whether a clean checkout can still regenerate the dashboard after the project-manifest rename.
Evidence. A 10,000-resample stratified bootstrap on the fixed transfer predictions gives:
| Cohort | Transfer AUC | 95% CI | Relation to chance |
|---|---|---|---|
| Portugal | 0.698 | 0.544 to 0.837 | above |
| India/Mumbai | 0.621 | 0.455 to 0.782 | uncertain |
| Mexico | 0.564 | 0.441 to 0.684 | uncertain |
| South Africa | 0.320 | 0.200 to 0.448 | below |
The South African reversal is statistically supported on this test set, while India and
Mexico cannot be separated from chance. The ordering does not establish a distance law: the
project never defined or tested a geographic, genetic, dietary, or microbiome population-
distance variable. One older entry also conflated two models: reversing pooled predictions at
AUC 0.320 yields 0.680, not the separately fit within-cohort CV AUC 0.853. Finally, dashboard
regeneration failed in a clean workbench because the generator still opened the renamed
polaris.toml and its three small extension count tables remained gitignored; its displayed
reproduction status was hard-coded to pass. The preprint dataset section also retained a
49-sample/175-total holdout count while the checked-in VSEARCH matrix and results use 50/176,
and it described the final training matrix as 10,353 samples (1,378 cases) while the artifact
used by the current smoke test contains 10,357 (1,380 cases). The four-sample difference is
confined to Shanghai; the older 10,353-sample analyses remain valid as pre-VSEARCH results.
Change. The data contract now computes deterministic transfer intervals and derives its
pass/check status from the stated tolerance. The dashboard shows interval whiskers, prints the
intervals in tooltips and the cohort table, and calls a cohort above/below chance only when its
full interval clears 0.5. Current-facing copy says transfer varies by cohort rather than
claiming an unmeasured distance gradient. The generator reads commandem.toml, and the three
small public-derived extension count tables are tracked so fresh-clone regeneration is real.
The preprint now reports the intervals, corrects the final matrix to 10,357 training samples
and the holdouts to 50/176, labels the 10,353-sample LOSO and model-comparison results as
pre-VSEARCH, separates the training permutation test from holdout evidence, and distinguishes
flipped pooled predictions (0.680) from the within-cohort model (0.853). Older log entries
remain intact with inline correction notes.
Confidence. High on the arithmetic, sample counts, generator failures, and conditional test-set intervals. The intervals condition on this one study sample and do not capture study- selection, label, preprocessing, or model-development uncertainty; the absence of a tested distance relationship is a scope correction, not evidence that distance has no effect.
2026-09-06. Adversarial audit: reproduction confirmed, two mechanism claims downgraded#
Trigger. An independent adversarial audit of the full research chain (experiments #1-#8, preprint, data contract), run before any release decision. Older entries are NOT rewritten; they carry short correction notes pointing here, so the record stays honest about what we believed when.
What the audit VERIFIED. All fast experiments reproduce exactly (re-run produced zero diff against the committed JSONs; baseline CV 0.81 / Portugal 0.698 / 3-cohort 0.673). The 71% flip prevalence recomputes by hand (57 of 80). The LOCO de-inversion, the sign-adaptation curve and its 0/4/5/7/10 flip counts, the OOD numbers, diversity-transfers-worse, and the 6x prevalence trend all check out. The 0.853 / 0.859 / 0.79 within-SA values are consistent (different models, each labeled correctly).
Finding 1 (downgrade). The "top-20 markers flip" table does not test the model's top markers.
experiment_marker_stability.py selects the top 20 by |mean point-biserial r| in the pooled
training populations. By that metric Atopobium ranks 133 of 139 (NHANES +0.02, China -0.04),
so the marker the narrative is about is not in the table. CatBoost importance (where Atopobium
is rank 1) and linear association strength are different rankings; the entries conflated them.
The table is also non-monotone: India and South Africa share 40% flips at transfer 0.621 vs
0.320. The flip-fraction-tracks-transfer claim is SUGGESTIVE for association-top markers, and
is no longer presented as the Atopobium mechanism.
Finding 2 (downgrade). Atopobium's "training direction -0.401" is a pooling artifact. The -0.401 was computed on the combined NHANES+China pool, whose composition confounds it: the Chinese studies are ~82% T2D (clinics) and NHANES ~9% (survey), so a genus that differs by study/specimen inherits a spurious pooled association with the label (Simpson's paradox). Within each study Atopobium's association is near zero (NHANES +0.02, China -0.04), and experiment #8 later showed it flips between two Chinese studies (+0.19 vs -0.46). There is no settled biological "training direction" for Atopobium; the within-SA +0.47 stands, but the -0.40 to +0.47 "flip" compared an artifact to a signal. This strengthens, not weakens, the experiment #8 conclusion that the instability is study-level.
Finding 3 (disclosure). The 61% vs 71% comparison has nested structure. The four Chinese sub-studies appear both as the within-China groups and, pooled, inside the cross-population set, so the two estimates are not independent, and their denominators differ (99 vs 79 genera informative in >=2 groups). The direction of the conclusion is unchanged; the comparison is descriptive, not a clean decomposition.
Finding 4 (precision). "~2.4x farther" for South Africa's Mahalanobis distance was the most favorable pairwise ratio. The range against the other cohorts is 1.96x (vs Mexico) to 2.44x (vs India), ~2.2x on average. Also, the domain classifier range is 0.986 to 1.000, so "0.99+" was imprecise for Portugal (0.986). And in sign_adaptation.json Portugal carries reversed=true because the LOGISTIC probe's k=0 transfer is 0.483; that flag describes the probe, not the flagship CatBoost (0.698), and should not be read as "Portugal reverses."
Net effect on the story. The core chain holds: the SA reversal is real (not depth), 71% of informative genera flip, restriction de-inverts safely, a few labels re-orient, detection is partly label-free, and the instability is mostly study-level. What is retired is the specific "Atopobium's training direction flips" mechanism; the honest mechanism evidence is the broad association anti-correlation (r = -0.237) plus Atopobium's high model importance and its within-China instability. Preprint sections 6, 6.1, 6.4, and 6.7 were revised accordingly.
2026-09-02. Experiment #8: population biology or study effect? Mostly study#
Trigger. Every prior experiment brackets the South Africa reversal but none says whether it is African population biology or a study-specific artifact. A second African cohort would settle it but none is public. Desk-feasible proxy: the training pool holds four Chinese sub-studies (anhui, shandong, sichuan, shanghai), same population, different labs and runs. If directions flip among same-population studies about as much as across populations, then study/batch alone reverses signs and the population-biology reading weakens.
Method. Per-genus point-biserial T2D association within each group. Two comparisons, plus an
Atopobium spotlight: association-vector correlation (as in the 2026-08-27 train-vs-SA analysis),
mean over within-China study pairs vs cross-population pairs; and sign-flip rate among genera
informative (|r| >= 0.15) in >= 2 groups. scripts/experiment_population_vs_study.py.
Result.
- Sign-flip rate: 61% within China (of 99 genera, same population, 4 studies) vs 71% across populations (of 79). Same-population studies flip nearly as often as entirely different ones. [Correction 2026-09-06: the comparison is nested, the four Chinese sub-studies also appear, pooled, inside the cross-population set, and the denominators differ (99 vs 79 genera), so the two rates are descriptive rather than an independent decomposition. Direction unchanged. See the audit entry.]
- Association-vector correlation: +0.163 within-China vs -0.001 cross-population. Population is not nothing (crossing populations collapses the weak agreement to zero), but same-population concordance is already weak.
- Atopobium flips within China too: anhui +0.19, shandong -0.46, sichuan -0.25, shanghai -0.14, a 0.65 swing inside one population, comparable to its cross-population range (SA +0.47).
Interpretation. The direction instability is mostly a study/batch phenomenon, not primarily population biology. Same-population studies already disagree on ~61% of directions; crossing populations adds only ~10 points. So the South Africa reversal is not distinguishable from generic study-level instability, and the tempting "African biology" explanation is not supported by public data. A weak population signal does survive (within-China correlation +0.16 exceeds cross-population 0.00), so population contributes a real but secondary increment.
Confidence. The within-vs-cross gap is consistent across both metrics and the Atopobium within-China flip is unambiguous. GRADED: the Chinese sub-studies are small (n=45 to 347) so their per-genus signs are themselves noisy; some within-China flipping is small-sample noise rather than batch per se, but that unreliability is exactly the finding, directional associations do not reproduce at these sample sizes, within or across populations.
Implication. Reframes the flagship. The transfer gap is dominated by non-reproducible, study-level direction estimates rather than clean population biology, which raises the bar for any biomarker claim and makes harmonized, multi-study-per-population data (not just more populations) the real prerequisite. It also means the SA reversal should be presented as a study-level effect first, with population as a weaker contributor.
2026-09-01. Experiment #7: going coarse does not help, community diversity transfers worse than taxa#
Trigger. Experiments #2-#6 show the taxonomic signal is real but direction-unstable across populations. Maybe a coarser summary, alpha diversity, is weaker but more portable, because a single scalar is less sensitive to which specific genera flip.
Method. Over the 139 model genera (a shared reference across cohorts) compute per-sample
Shannon, Simpson, and richness from raw counts. A logistic model on the three diversity features
is trained on the pooled cohorts, cross-validated within each cohort, and transferred to each
unseen cohort, then compared to the 139-genus taxonomic model. Training pool and Portugal come
from the unified raw-count table joined to the pkl sample-id labels; India/Mexico/SA from their
interim counts. scripts/experiment_diversity_transfer.py.
Result.
| Cohort | diversity within-CV | diversity transfer | taxa transfer | Shannon (T2D - ctrl) |
|---|---|---|---|---|
| Portugal | 0.27 | 0.432 | 0.698 | +0.068 |
| India/Mumbai | 0.50 | 0.544 | 0.621 | -0.093 |
| Mexico | 0.62 | 0.451 | 0.564 | +0.156 |
| South Africa | 0.79 | 0.188 | 0.320 | +0.056 |
- Diversity transfers WORSE than taxa (mean 0.404 vs 0.551) and reverses in 3 of 4 cohorts (taxa reverse in 1). South Africa reverses even harder under diversity (0.188 vs 0.320).
- The direction of diversity itself is population-specific: Shannon is higher in T2D in most cohorts but lower in India, so the sign flips like the taxa do.
- Diversity does carry a within-population signal (training CV 0.639; South Africa within 0.79), it just does not travel.
Interpretation. Coarse-graining is not a shortcut to portability. The direction instability is baked into the community-disease relationship at every level, right down to a single number: collapsing to diversity throws away the specific taxonomic signal while keeping the instability. This refutes the "coarse features are more portable" hypothesis.
Confidence. Clean comparison on shared features. GRADED DOWN: within-CV AUC for n=50 cohorts is noisy (Portugal 0.27 is a small-sample artifact, not a real anti-signal); the reliable read is the transfer comparison, where diversity is clearly worse and reverses more.
Implication. No free lunch from summary statistics. Portability has to come from modeling direction (experiments #3, #4), not from a coarser feature space.
2026-09-01. Experiment #6: what makes a marker portable? Stable markers are the prevalent ones#
Trigger. Experiment #2 split the 139 genera into direction-stable markers and sign-reversing flippers but did not ask why. If flippers are just the rare, low-abundance taxa whose association sign is noisy, a portable panel should favor prevalent, abundant genera, and the prevalent-yet- reversing genera are the interesting exceptions.
Method. For each genus compute two label-free properties from the full raw count table:
prevalence (fraction of samples detected) and mean relative abundance. Join with the stability
class from experiment #2; compare stable markers vs flippers (Mann-Whitney), fit a small logistic
(flipper ~ log-prevalence + log-abundance), and flag the prevalent+abundant flippers.
scripts/experiment_marker_properties.py.
Result.
- Direction-stable markers are about 6x more prevalent than flippers (median 21% of samples vs 3.6%); abundance shows the same direction. The flipper-vs-stable logistic loads most on low prevalence (standardized coef -0.34).
- But the difference is not statistically significant (prevalence p=0.13, abundance p=0.12), with only 15 stable genera, so it is a suggestive trend, not a proven rule.
- A distinct set of prevalent, abundant genera reverse anyway: Faecalibacterium, Veillonella, Akkermansia, Actinomyces, Alloprevotella, Atopobium, Campylobacter, and more. These reverse for reasons other than sparsity and are the biologically meaningful reversers.
Interpretation. Rarity explains part of the flipping (rare taxa have noisier association signs), so a portable panel should prefer prevalent, abundant genera. But prevalence is no guarantee: the dominant markers a model relies on (Atopobium) can be both common and reversing, which is exactly why the pooled model fails confidently rather than quietly.
Confidence. The prevalence gap is a clear 6x effect but underpowered (p0.12-0.13); the
prevalent-flipper list is descriptive. Properties are computed on the training-pool raw counts.
Implication. Actionable for panel design: start from prevalent, abundant genera, but test each marker's direction per population rather than trusting prevalence alone.
2026-09-01. Experiment #5: can the model flag its own failure, without labels?#
Trigger. Experiment #4 re-orients a reversed population but needs local labels. The open question is detection rather than correction: before any labels, can a deployment tell it is out-of-distribution (OOD) and likely to answer backwards, and refuse instead of being confidently wrong?
Method. Reference = the pooled training distribution in the 139-genus rank space. Two
label-free OOD scores per unseen cohort: (1) a domain-classifier AUC (5-fold CV AUC separating
training from the cohort, balanced downsample, 15 seeds), and (2) the Mahalanobis distance of
the cohort centroid from training (Ledoit-Wolf shrinkage covariance). Then ask whether either
score, computed with no diabetes labels, flags the reversed cohort. scripts/experiment_ood_detection.py.
Result.
| Cohort | Transfer | Domain-classifier AUC | Mahalanobis |
|---|---|---|---|
| South Africa (reversed) | 0.320 | 1.000 | 20.3 |
| Mexico | 0.564 | 0.998 | 10.3 |
| India/Mumbai | 0.621 | 0.996 | 8.3 |
| Portugal | 0.698 | 0.986 | 8.8 |
- The domain classifier saturates: every cohort is ~perfectly distinguishable from training (0.986 to 1.000), because batch and study effects make any new cohort trivially OOD. Its ranking is perfectly monotonic with transfer (Spearman -1.0 over the four cohorts) but the spread is only 0.014, so a threshold would refuse everyone, including Portugal (transfers fine).
- Mahalanobis distance separates cleanly: South Africa's centroid sits ~2.4x farther out (20.3) than any other cohort (8 to 10). The reversed cohort is the distributional outlier. [Correction 2026-09-06: 2.4x was the most favorable pairwise ratio; the range is 1.96x (vs Mexico) to 2.44x (vs India), ~2.2x on average. See the audit entry.]
Interpretation. Detection of a gross reversal is partly possible without labels: a distance threshold flags South Africa for refusal, so a deployment need not answer backwards. But two honest limits remain. A domain classifier alone cannot triage (batch effects make everything look OOD), and even the usable distance signal cannot rank the non-reversed cohorts by transfer quality or distinguish attenuation from reversal among them. Being different is only loosely tied to being wrong.
Confidence. Four unseen cohorts, so any OOD-versus-transfer correlation is weak and the Spearman -1.0 rests on tiny domain-AUC gaps; the robust part is the Mahalanobis outlier (South Africa 2.4x the rest), which is a clear, single-number effect. The claim is deliberately modest: partial unsupervised detection, not reliable triage.
Implication. A safe deployment can pair a label-free distance gate (refuse or down-weight grossly OOD populations) with the label-based calibration of experiment #4 (re-orient once a few local labels arrive). Detection buys caution; correction still buys accuracy.
2026-08-28. Experiment #4: per-population sign adaptation, a few local labels re-orient the reversed model#
Trigger. Experiment #3 de-inverted South Africa by dropping the sign-reversing genera, but that throws away real (if flipped) signal and costs accuracy. The better move is to ADAPT the marker directions to the target population rather than discard them. Pure unsupervised sign correction is impossible (you cannot know a population is flipped without labels), so the honest question is few-shot: with k labeled local samples, can we re-orient the global model and recover the within-cohort signal it reads backwards?
Method. scripts/experiment_sign_adaptation.py. Interpretable logistic regression in the
139-genus rank space, where a coefficient's sign is a marker's diabetes direction. A global model
is fit on the NHANES+China pool (beta_global). For a target cohort we fit a prior-regularized
logistic on k local samples: minimize logloss + (lambda/2)||beta - beta_global||^2, so at k=0 it is
the global model and as k grows coefficients can flip sign where the population demands it, shrunk
toward the global prior for stability. Leakage guard: within each cohort, calibration and test
samples are disjoint (repeated stratified 5-fold, 12 seeds); the global model never saw the cohort.
We sweep k in {0, 5, 10, 20, 40} and count coefficient sign flips vs the global model.
Result (AUC by number of local labels):
| Cohort | k=0 | k=10 | k=20 | k=40 | within-cohort ceiling |
|---|---|---|---|---|---|
| South Africa | 0.320 | 0.405 | 0.487 | 0.629 | 0.859 |
| Portugal | 0.483 | 0.511 | 0.517 | 0.562 | 0.766 |
| India/Mumbai | 0.545 | 0.557 | 0.573 | 0.587 | 0.639 |
| Mexico | 0.620 | 0.619 | 0.626 | 0.633 | 0.627 |
- South Africa, read backwards at k=0 (0.320), climbs monotonically and crosses chance by ~40 local labels (0.629), heading toward its within-cohort ceiling of 0.859.
- The mechanism is visible: South Africa's coefficient sign flips vs the global model grow 0, 4, 5, 7, 10 across the k grid. Adaptation is literally re-orientation of markers.
- Cohorts that were not reversed gain little (Mexico has no signal to recover; India and Portugal inch up), which is the expected pattern.
Interpretation. The reversal is fixable, but only with local supervision. A small calibration set re-orients the flipped markers and turns a confidently-wrong model into a useful one, which is exactly "per-population sign adaptation." It also sets a rough cost: on the order of tens of labeled local samples to de-invert South Africa, not hundreds.
Confidence. The de-inversion trend and the growing sign-flip count are robust (leakage-free splits, 12 seeds). GRADED DOWN: this is an interpretable logistic probe, weaker than the flagship CatBoost, so its k=0 transfer values are not the leaderboard AUCs (Portugal's logistic transfer is 0.483 here versus 0.698 for CatBoost); the point is the mechanism and the trend, not the absolute level. Even at k=40, South Africa (0.629) is well short of its 0.859 ceiling, so tens of labels help but do not fully recover the signal, and per-cohort n is small.
Implication. The deployable design is a global model plus a lightweight local calibration step, not a single frozen classifier. Any future African/SCCS cohort gives a direct test: does the same handful of local labels re-orient it the same way?
2026-08-28. Experiment #3: portability over accuracy, a stable-feature model stops failing backwards#
Trigger. Experiment #2 showed the transfer gap is a direction problem: 71% of informative genera flip sign across populations and the model's top markers flip 40 to 50% in distant cohorts. The convicted hypothesis: a model restricted to the direction-stable genera should transfer more honestly (fewer confident reversals), even at the cost of raw accuracy, because it drops the flippers the pooled model reads backwards.
Method. scripts/experiment_portability.py. Two tests, both leakage-free.
- A. Stable-restricted vs full transfer. The stable set from experiment #2 was defined using the test cohorts, so restricting features and testing on those same cohorts would be circular. So we select leave-one-cohort-out: to evaluate transfer to cohort T, the direction-stable genera are chosen from the OTHER five cohorts only, so T's labels never inform its own feature set. Training always uses only the NHANES+China pool. Compare the stable-restricted CatBoost to the full 139-feature model on each held-out cohort.
- B. Mechanism ablation. Drop Atopobium, the top marker chosen by TRAINING importance alone (leakage-free), and measure South Africa transfer.
Result.
- Restriction trades accuracy for portability. Portugal falls 0.698 to 0.589 (the full model is better where transfer already works), India holds (0.621 to 0.616), Mexico edges up (0.564 to 0.600), and South Africa de-inverts: 0.320 to 0.515, no longer read backwards.
- Reversed cohorts: 1 to 0. Mean transfer rises modestly, 0.551 to 0.580.
- Ablation: dropping Atopobium alone moves South Africa 0.320 to 0.410, a partial fix, so the reversal is broad and not carried by a single marker (consistent with the 2026-08-27 finding that 7 of the top 12 genera flip).
Interpretation. You cannot build a more accurate global model this way, but you can build a safer one: restricting to direction-stable genera removes the confident reversal and makes the model fail to chance rather than fail backwards. This is the concrete payoff of "portability over accuracy," and it validates experiment #2's thesis with a leakage-free causal test: the flippers are what break transfer, and removing them fixes the break at a measurable accuracy cost.
Confidence. The de-inversion and the reversed-cohorts 1-to-0 result are robust to leakage (LOCO selection, training on the training pool only). GRADED DOWN: a move toward 0.5 is a loss of confident-wrongness, not a gain of skill; per-cohort n is small (50 to 88); and the mean-transfer gain (0.03) is modest and within noise for cohorts this size. The honest headline is safety, not accuracy.
Implication. The screening target should be a portability-first model (direction-stable features, or per-population sign adaptation), not a single global classifier optimized for AUC. Any future African/SCCS cohort tests whether the same stable set keeps South Africa non-reversed.
2026-08-28. Experiment #2: which genera generalize? Direction, not strength, fails to transfer#
Trigger. The South Africa reversal showed Atopobium flips its diabetes direction (control-associated in the training populations, case-associated in SA). That convicts a systematic question across ALL six populations: which oral genera keep a consistent T2D direction (trustworthy markers) and which flip? A model can only transfer on the stable ones.
Method. Six populations placed in the shared 139-genus space: NHANES (US) and China from
the training pool (split by study_labels), Portugal from the geographic holdout, and
India/Mumbai, Mexico, South Africa as external cohorts, each per-sample rank-transformed
exactly as at inference (leakage-free). Per genus per cohort: point-biserial correlation
between the sample's rank for that genus and the T2D label (sign = direction, |r| = strength).
Stability is defined by DIRECTION consistency, not magnitude: NHANES (n=9683) has a diffuse
signal (median |r| = 0.01), so a fixed magnitude gate would wrongly mark the largest cohort
uninformative. A cohort casts a sign vote once |r| >= 0.05; a genus is "robust" when its sign
agrees across >= 5 of 6 cohorts, a "flipper" when it clears |r| >= 0.15 in both directions.
scripts/experiment_marker_stability.py.
Result.
- 71% of genera informative (|r| >= 0.15) in >= 2 populations reverse sign across them (57 of 80). This figure excludes NHANES, so it does not depend on the magnitude artifact.
- Only 15 genera hold a single direction across >= 5 populations, and all are weak (|mean r| 0.03 to 0.12). Strongest stable case-associated: Streptococcus (+0.12), Peptoniphilus, Enterococcus, Bifidobacterium. Stable control-associated: Fretibacterium, Johnsonella, Treponema, Cloacibacterium.
- 56 flippers, Atopobium among them (NHANES +0.02, China -0.04, Portugal -0.07, India +0.15, Mexico -0.19, SA +0.47). Others: Faecalibacterium, Campylobacter, Bacteroides, Actinomyces.
- Effect: marker direction predicts transfer. Of the pooled model's 20 strongest training markers, the fraction that flip sign tracks how well the model transfers: Portugal 10% (transfer AUC 0.698), India 40% (0.621), Mexico 50% (0.564), South Africa 40% with the highest-importance marker among them (reversed to 0.320). [Correction 2026-09-06: "strongest training markers" here means strongest by pooled ASSOCIATION, not by model importance; Atopobium ranks 133/139 by that metric and is not in this set, and the relation is non-monotone (India = SA = 40%). See the audit entry.]
Interpretation. The cross-population transfer gap is a direction problem, not a strength problem. The diabetes signal is real and often strong inside each population, but the SIGN of almost every genus-to-diabetes association is population-specific. South Africa is not an outlier so much as the extreme of a pervasive instability. This is why a pooled model degrades smoothly with population distance and then breaks entirely on SA: it is averaging directions that do not agree. The few direction-stable genera are too weak to carry a portable model on their own.
[Correction 2026-09-25: no population-distance variable was defined or tested. The supported claim is that transfer varies sharply by cohort, with a statistically supported reversal in South Africa; a smooth distance relationship is retired. See the audit entry.]
Confidence. The 71% reversal prevalence and the marker-flip-vs-transfer effect are robust to the NHANES magnitude artifact (the effect uses only sign, the prevalence excludes NHANES). GRADED DOWN: these are associations in modest per-cohort samples (n = 50 to 88 for four of six cohorts), not causal or diagnostic claims, and several flippers are low-abundance or environmental/reagent-associated taxa (Serratia, Burkholderia, Delftia, Paenibacillus) whose signs may reflect contamination rather than biology. The "robust" panel is direction-consistent but weak, not a validated biomarker set.
Implication. Points the modeling work at portability rather than accuracy: a model built on direction-stable genera, or one that adapts sign per population, is the honest path to transfer, and any future African/SCCS cohort gives a direct test of whether the flippers (Atopobium first) flip the same way. It also reframes the flagship claim: the oral-microbiome T2D signal is real but locally oriented, so a single global screening model is the wrong target.
2026-08-27. The SA reversal is broad and led by Atopobium (the model's top marker)#
Trigger. South Africa reverses (strong within-cohort signal, backwards transfer). Which genera flip, and is it a few markers or a global inversion?
Method. For each of the 139 model genera, point-biserial correlation of its
(rank-transformed) abundance with the T2D label, computed in the TRAINING populations and
in South Africa; then correlate the two per-feature association vectors and inspect the
model's most important features. scripts/experiment_feature_reversal.py.
Result.
- Per-feature T2D associations, training vs SA: r = -0.237 (p = 0.005), a significant broad anti-correlation.
- The model's dominant marker Atopobium (importance ~19, several times the next) flips hard: assoc_train -0.401 (lower in diabetics) to assoc_SA +0.467 (higher in diabetics). [Correction 2026-09-06: the -0.401 is a POOLED NHANES+China estimate confounded by study composition (China ~82% T2D vs NHANES ~9%); within each study Atopobium's association is near zero (+0.02 / -0.04). The within-SA +0.47 stands; the "training direction" does not. See the audit entry and experiment #8.]
- 7 of the top 12 model genera flip diabetes-direction in SA (Atopobium, Phocaeicola, Alloprevotella, Treponema, Neisseria, Mogibacterium, Rhodobacter).
Interpretation. The reversal is broad (significant negative association-correlation) and concentrated in the genera the model relies on most, led by Atopobium. Because Atopobium alone carries ~19% importance and its diabetes association strongly flips, the pooled model is confidently backwards on SA. It is not a single-feature artifact: most of the model's reasoning inverts. Notably, Atopobium is the preprint's headline T2D marker, so its reversal in South Africa is exactly the finding to scrutinize.
Confidence. The Atopobium flip and the significant negative correlation are clear. Whether this reflects population biology or a SA-study batch effect is still unresolved, a systematic case/control batch difference could also flip associations broadly.
Implication. Names the mechanism (Atopobium-led broad reversal) and gives a concrete, pre-registered thing to check in any future African or SCCS cohort: does Atopobium's diabetes direction flip there too? If it does, population biology gains support; if not, the SA study is idiosyncratic.
2026-08-27. No second public African oral cohort to replicate the reversal#
Trigger. South Africa shows a strong, reversed T2D signal. To decide whether that is population biology or specific to one study, we need an independent African (or African-descent) oral 16S cohort with diabetes labels.
Method. Systematic search across NCBI SRA/BioProject, ENA, GEO, Qiita, PubMed, Google Scholar, and African-microbiome/H3Africa angles (2019-2026), filtering for oral (not gut) 16S with per-subject T2D or glycemic labels.
Result. No clean match exists. Public African oral 16S with T2D labels is essentially a single cohort, the one we already have (PRJNA723337). Nearest options, each caveated:
- Uganda PRJNA1087153 (public): gingival crevicular fluid, Nanopore full-length 16S, n=45 (26 diabetes / 19 non-diabetes), diabetes diagnosis but no HbA1c. Different specimen and platform, so only a weak directional check, not a like-for-like replication.
- SCCS (African-American, US): mouth rinse, Illumina V4, n=294, real T2D labels, the methodologically closest resource, but CONTROLLED ACCESS (requires a data-use request).
- Ruled out: Nigeria (samples pooled, no deposition), Egypt PRJNA1475433 (healthy-only), Angola/Zimbabwe saliva (no diabetes labels). A "SCCS on SRA" lead was a mislabel, the accession is a rhesus-macaque project.
Interpretation. The population-vs-study question CANNOT be settled with off-the-shelf public data today. This is a genuine data-availability limitation, not a failure to search: African oral-microbiome + diabetes data is thin (most African diabetes-microbiome work is gut/stool).
Confidence. High on the availability conclusion (convergent search across repositories).
Implication. Paths forward: (a) run the Uganda cohort as an explicitly weak, cross-platform directional check; (b) apply for SCCS controlled access, the rigorous answer; (c) partner for new data (e.g. H3Africa-linked oral cohort). Until then, the SA reversal is reported as a real, verified, single-cohort finding whose population-generality is unresolved.
2026-08-27. South Africa's inversion is a REAL, reversed signal (not noise)#
Trigger. Depth was ruled out. Is SA's inversion (AUC 0.320) a genuine reversed signal, or just small-sample noise that happens to land below 0.5?
Method. For each new cohort, in the SAME 139-feature rank space, compare the
TRANSFER AUC (pooled model scored on the cohort) with a WITHIN-COHORT AUC (5-fold CV
of a model trained and tested inside the cohort). scripts/experiment_within_cohort.py.
Result.
| Cohort | n | Transfer AUC | Within-cohort CV |
|---|---|---|---|
| India/Mumbai | 60 | 0.621 | 0.638 |
| Mexico | 87 | 0.564 | 0.566 |
| South Africa | 88 | 0.320 | 0.853 |
Interpretation. Three distinct regimes:
- South Africa has a strong, learnable T2D signal (within-CV 0.853) that is reversed relative to the training populations, which is exactly why the pooled model lands at 0.320. Flip the predictions in SA and the same model would score ~0.85. [Correction 2026-09-25: flipping these pooled-model scores gives AUC 0.680 (= 1 - 0.320), not 0.85. The 0.853 figure is a separately trained within-cohort CV model; it shows learnable local structure but is not the flipped pooled model. See the audit entry.] So the inversion is real signal pointing the opposite way, not noise and not depth.
- Mexico is ~chance both ways (0.566 within, 0.564 transfer): little learnable T2D signal at all, consistent with its 100%-periodontitis composition masking the contrast.
- India transfers about as well as it is internally learnable (0.638 vs 0.621).
So "cross-population failure" is not one thing: it has at least two modes, no signal (Mexico) and reversed signal (South Africa).
Confidence. The SA reversal is clear: within-CV 0.853 is far above chance, so even allowing for small-n CV optimism there is strong, real signal that the pooled model gets backwards. What one African cohort CANNOT separate: population-biological reversal vs a study/batch-specific reversal (SA is a single study).
Implication. The oral-microbiome-to-T2D relationship can genuinely differ, even reverse, across populations. A pooled model is not merely unreliable across populations, it can be confidently wrong, which argues for population-specific calibration (or sign checks) before any deployment. Next: a second African (or African-descent) cohort to test whether the reversal reproduces (population) or is specific to this study.
2026-08-27. South Africa inversion is NOT a sequencing-depth artifact#
Trigger. SA transferred at AUC 0.320 (inverted). The leading suspect was sequencing depth: SA is deep (~218k mapped reads/sample) vs the shallower cohorts, and deep sequencing detects more rare genera → a denser per-sample rank profile that could be out-of-distribution for the training data.
Method. Rarefied each SA sample's genus-count vector to a ladder of shallower
depths (50k → 1k, multinomial subsample, 3 seeds averaged), re-applied the identical
per-sample rank transform, and re-scored with the pooled-cohort CatBoost.
scripts/experiment_sa_rarefaction.py.
Result.
| SA depth | Transfer AUC |
|---|---|
| full (~218k) | 0.320 |
| 50,000 | 0.358 |
| 20,000 | 0.406 |
| 10,000 | 0.376 |
| 5,000 | 0.380 |
| 2,000 | 0.369 |
| 1,000 | 0.377 |
Total AUC swing across depth: 0.048, and it stays inverted (well below 0.5) at every depth, down to 1,000 reads.
Interpretation (hypothesis rejected). Sequencing depth is not the driver of the inversion. Rarefying to matched-or-shallower depth nudges AUC up by at most ~0.05 but never uncrosses 0.5. So the model's T2D-associated oral-microbiome features genuinely anti-correlate with T2D in this African cohort, a real cross-population reversal, not a technical/depth artifact.
Confidence. Solid rejection of the depth hypothesis (consistent across six depths × three seeds). What remains open is why it inverts, population biology vs a study/batch confound vs cohort composition, which one deep cohort can't separate.
Implication. Strengthens the core thesis: cross-population failure is real and can be directional, not just noisy. Next suspects to test: study/batch effect (SA is a single-study cohort) and whether a second African/saliva cohort reproduces the sign.
2026-08-23. South Africa (PRJNA723337): signal INVERSION (AUC 0.320, verified)#
Trigger. Third plaque cohort, completing the geographic sweep (Africa gap, the only public African oral 16S T2DM data).
Method. PRJNA723337, 88 subgingival-plaque samples (57 diabetic [DM+Known_DM] /
31 control [Normal]; Pre-DM excluded). Single-end, deep (~227k reads/sample).
Required --fastq_qmax 50 (SA reads hit Q43, above vsearch's default 41, a real
pipeline fix, now in the generic processor). Labels verified against NCBI BioSample
isolate (spot-checked: DM_13→diabetic, Known_DM_23→diabetic, etc., all correct).
Same 139-feature → rank → CatBoost(existing pool) transfer test.
Result (verified inversion).
| value | |
|---|---|
| Transfer AUC | 0.320 (below random) |
| Flipped-label AUC | 0.680 (systematic, not noise) |
| Mean pred: diabetic vs control | 0.829 vs 0.884 (controls scored more diabetic) |
| Genera observed | 125/139 (highest, deep sequencing) |
Four-cohort gradient (all new cohorts this session):
| Population | Specimen | Depth | Transfer AUC |
|---|---|---|---|
| European (Portugal) | saliva | n/a | 0.698 |
| S. Asian (India) | plaque | ~97k | 0.621 |
| Latin Am. (Mexico) | saliva | ~50k | 0.564 |
| African (South Africa) | plaque | ~227k | 0.320 (inverted) |
Interpretation. A real, verified systematic inversion, the extreme end of cross-population failure: the model's T2DM-associated features anti-correlate with T2DM in South Africa. The gradient above roughly tracks population distance from the US+Asian training set. BUT the cause is entangled and must not be overclaimed as biological: SA differs on four axes at once, population (African), specimen (plaque), sequencing depth (227k reads → far denser genus detection → a very different rank-profile distribution than the shallow training cohorts), and clinic composition. The compressed, uniformly-high predictions (0.83–0.88) point to a severe distribution shift where the rank transform interacts badly with the depth difference.
[Correction 2026-09-25: the apparent ordering is descriptive; no population-distance measure or distance-transfer test was specified. Current-facing surfaces no longer call it a gradient.]
Confidence. Inversion is verified and systematic (labels correct, flip=0.68). Attribution is NOT established, do not claim "African oral T2DM signal is inverted."
Implication. (1) Naive cross-population application isn't just unreliable, it can be actively wrong, a strong argument for mandatory population-specific calibration before any deployment. (2) New concrete suspect: sequencing-depth harmonization. Testable next step, rarefy/normalize SA to comparable depth before the rank transform and re-run; if the inversion softens, depth (not population) is a major driver.
2026-08-23. Mexico (PRJNA1208116) transfer test REFUTES the specimen hypothesis#
Trigger. India (plaque) transferred only partially and naive pooling hurt; the leading hypothesis was that specimen mismatch (plaque vs saliva) was the culprit. Mexico is saliva, a specimen-matched, balanced, new-continent test to isolate population effect from the plaque confound.
Method. PRJNA1208116, 87 saliva samples (41 T2DM / 46 control; labels recovered
from the paper's Supplementary Table S1, CC-BY). Same VSEARCH → 139-feature → per-sample
rank → CatBoost(existing pool) transfer test. scripts/integrate_cohort.py.
Result.
| Holdout | Specimen | Population | AUC |
|---|---|---|---|
| Portugal | saliva | European | 0.698 |
| 3-cohort | mixed | mixed | 0.673 |
| India | plaque | S. Asian | 0.621 |
| Mexico | saliva | Latin American | 0.564 |
| Asian↔US (known) | n/a | n/a | ~0.541 |
Mexico had the best feature overlap (109/139) yet transferred worst, near random.
Interpretation (hypothesis refuted). Specimen match did NOT rescue transfer: saliva Mexico (0.564) transferred worse than plaque India (0.621). So specimen type is not the dominant barrier. BUT the result is confounded: Mexico is a 100%-periodontitis cohort (Severe 48 / Moderate 31 / Mild 21), so the T2DM contrast is within periodontitis, where the diabetes signal is subtle relative to the dominant periodontitis-driven microbiome shift, and this differs from the general-population (NHANES-heavy) training set. Two entangled explanations remain: (a) population distance (Latin American), (b) cohort composition (all-periodontitis clinic case-control vs general-population survey). One cohort cannot separate them.
Confidence. Result solid (near-random, n=87, balanced). Interpretation graded: specimen-hypothesis refutation is well-supported; population-vs-periodontitis remains open.
Implication (reframes the roadmap). The clinic-recruited, periodontitis-focused cohorts (India, Mexico) test a different, harder, confounded task than general-population T2DM screening. To measure pure population transfer, prioritize general-population survey cohorts (like NHANES itself, or the Qatar Biobank) over periodontitis clinic case-controls. This is now the sharper dataset-search criterion.
2026-08-23. Does pooling India into training help the existing holdouts? (NO)#
Trigger. After the India transfer test, ask the complementary question: does adding India to the training pool improve generalization to the current holdouts?
Method. Baseline CatBoost (existing pooled cohorts) vs. the same with India's 60
plaque samples appended to training (rank-transformed into the 139-feature space).
Re-scored the Portugal and 3-cohort holdouts. scripts/experiment_india_in_pool.py.
Result (negative).
| Model | Portugal | 3-cohort |
|---|---|---|
| baseline (no India) | 0.698 | 0.673 |
| + India pooled | 0.624 (−0.074) | 0.621 (−0.052) |
Interpretation. Naive pooling of the India plaque cohort degrades transfer to the (saliva) holdouts by 0.05–0.07 AUC. The specimen mismatch (plaque vs saliva) plus partial feature overlap acts as a domain shift that adds noise rather than useful diversity. Consistent with the project's recurring finding that naive cross-study pooling hurts generalization (prior domain-adaptation attempts also degraded holdout).
Confidence. Clear directional result, both holdouts drop materially and in the same direction. Caveat: single cohort, n=60, specimen-confounded.
Implication. Don't naively pool across specimen types. India (plaque) needs site-matching or explicit domain handling; specimen type is a confound to control before added geography can help. Motivates seeking a saliva cohort for new geography.
2026-08-23. India/Mumbai (PRJNA1240053) cross-population transfer test#
Trigger. Living-project goal "close the cross-population transfer gap": test whether the existing pooled model generalizes to a never-seen population.
Method. Downloaded the India/Mumbai plaque 16S cohort (60 samples, labels from
SRA aliases: 40 diabetic [T2DM + T2DM+perio] / 20 non-diabetic [healthy + perio]).
Processed through the same VSEARCH pipeline (merge → maxee 1.0/minlen 200 filter →
97% closed-ref vs HOMD v16.03 → genus). Mapped to the 139 model genera, applied the
identical leakage-free per-sample rank transform, and scored with CatBoost
(depth 4, iters 200, lr 0.1, balanced, seed 42) trained on the existing pooled
cohorts. India was never seen in training, a clean LOSO-style holdout.
Scripts: scripts/process_india_vsearch.py, scripts/integrate_india.py.
Result.
| Holdout | AUC | n |
|---|---|---|
| India/Mumbai (new) | 0.621 | 60 |
| Portugal (ref) | 0.698 | 50 |
| 3-cohort combined (ref) | 0.673 | 176 |
| Asian↔US (known failure) | ~0.541 | n/a |
Only 76/139 model genera were observed in India (plaque community + region differences); the 63 absent genera tie at the bottom rank and add little signal.
Interpretation. Partial transfer, above the near-random cross-continental failure floor (0.541) and above chance, but below the European holdouts. The model retains some signal on a new population, which is notable given three stacked disadvantages: (1) never-seen population, (2) plaque, not saliva (specimen mismatch), (3) only 55% feature overlap. The plaque mismatch likely understates population transfer, so 0.621 is a conservative floor for "does the signal reach India."
Confidence. Suggestive, not definitive, n=60, single cohort, specimen mismatch. Not strong enough to claim India generalization; strong enough to justify adding saliva Indian data and testing whether site-matching lifts it.
Next. (a) Get a saliva Indian cohort to separate site effect from population effect. (b) Add India to the training pool and re-check whether it helps other holdouts (diversity gain). (c) Pending Mexico/South Africa labels for more geography.