Plos iconPlosSep 23, 2026 ~6 min source read

Subsampling strategies for pairwise analysis of large pathogen genomic and spatial datasets: practical comparison using Mycobacterium tuberculosis data

A direct comparison of two subsampling approaches shows that a pairwise case-control design with five controls per case reduces computational cost while preserving unbiased inference and credible interval coverage for estimating associations with transmission clustering.

Comparing subsampling strategies for efficient pairwise analysis of large pathogen genomic and spatial datasets: An application to <i>Mycobacterium tuberculosis</i> tra...

Share this story

Send the public story page.

Useful takeaways from this story.

Case-control subsampling (5 controls per case) produced negligible bias and valid 95% credible interval coverage across simulation settings.

Divide-and-conquer Bayesian fitting performed comparably on average but showed wider bias dispersion and more outliers when true effect sizes were large.

Analysis used a real, spatially referenced genomic dataset of 4,154 Mycobacterium tuberculosis isolates and simulation studies representing three genomic clustering scenarios.

# Why this matters Pairwise analyses that relate genomic similarity and spatial or host covariates can estimate who transmitted infection to whom. Those analyses require evaluating every possible pair of individuals (dyads), which grows quadratically with sample size and quickly becomes computationally infeasible for moderately large datasets. This study compares two subsampling strategies intended to keep analyses tractable while preserving valid inference.

# Data and goal

# Methods compared

  • Divide-and-conquer Bayesian model fitting: partition the full dataset into subsets, fit models separately, then combine results. This aims to parallelize and reduce per-fit complexity but requires a valid strategy to aggregate estimators.
  • Pairwise case-control approach: treat dyads where isolates share genomic cluster membership as cases and sample a fixed number of control dyads (non-clustered pairs) per case. The study evaluated using five controls per case as a default.

The comparison used simulation studies based on three representative tuberculosis genomic clustering scenarios to assess bias, credible interval coverage, and distribution of estimates.

# Main findings Across all simulation scenarios, the case-control subsampling approach delivered negligible bias and achieved the expected 95% credible interval coverage for estimated associations. Its performance matched or exceeded the divide-and-conquer approach on key inferential metrics.

When true effect sizes were large, the case-control approach produced a somewhat narrower distribution of bias estimates and fewer extreme outliers than divide-and-conquer. Overall, the case-control method was more stable across simulated conditions examined.

# Practical recommendation When analyzing large genomic-spatial datasets where full pairwise Bayesian modeling is not feasible, use the pairwise case-control subsampling strategy with five controls per case. This reduces computational burden while maintaining unbiased point estimates and valid credible intervals in the settings tested.

# Additional context and reproducibility

# What this does not claim

# Bottom line For pairwise association studies of transmission using large genomic and spatial datasets, a case-control subsampling approach with five controls per case provides a practical balance: it substantially reduces computational cost and preserves unbiased estimation and credible interval coverage under the evaluated scenarios.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app