Bioengineer iconBioengineerSep 12, 2026 ~6 min source read

How multi-omics models handle entire missing data layers to predict patient outcomes

A new review maps methods that let predictive models work when whole omics modalities—genome, transcriptome, methylome, proteome, metabolome—are absent for many patients, and explains why usual fixes can bias results or destroy power.

AI Learns to Predict Patient Outcomes Even When Key Omics Data Are Missing

Share this story

Send the public story page.

Useful takeaways from this story.

Block-wise modality missingness (entire omics layers absent for some patients) is common in clinical cohorts and breaks methods built for fully paired data.

Discarding patients with any missing layer or imputing individual values both create serious problems: loss of statistical power, sampling bias, and misleading confidence.

Emerging solutions fall into three families: missingness-aware fusion, shared latent representations with subset-conditioned inference, and modality-completion via generative models.

# Why missing omics layers matter

# Common workarounds and their costs

# A taxonomy of better approaches A review by Ricky Nguyen and Fatemeh Vafaee (University of New South Wales), published in Artificial Intelligence Review, organizes newer methods into three families that explicitly address block-wise missingness.

Missingness-aware fusion architectures

  • These models encode which modalities are present for each patient and adapt how they combine inputs. Instead of requiring every layer, they learn fusion functions that accept arbitrary subsets of omics. The model treats absence as a condition to reason about, letting it weight available views by both their information content and the presence pattern.

Shared latent representations with subset-conditioned inference

Modality-completion (generative surrogates)

  • Generative frameworks—adversarial networks and autoencoders among them—are trained on the complete subset to synthesise plausible surrogates for missing omics layers. The review emphasizes the practical goal: generate estimates that help downstream predictors, not to recreate the exact molecular profile of an unassayed patient.

# Practical implications for clinical research Treating missing omics as structured information preserves sample size and reduces selection bias compared with complete-case filtering. Methods that reason about which modalities are present avoid the false confidence of point-wise imputation. Choosing among architectures depends on the cohort, the types of omics involved, computational resources, and whether the priority is interpretability or predictive performance.

# What to watch for when reading studies Look for explicit descriptions of how models handle modality absence: do they encode presence patterns, learn shared latent spaces, or generate surrogates? Check whether authors compare to complete-case and imputation baselines and whether they discuss cohort shrinkage or potential sampling bias introduced by filtering. Methods that model missingness directly provide more credible claims when cohorts are heterogeneously assayed.

# Bottom line Block-wise missingness is a common, consequential pattern in multi-omics clinical datasets. The growing methodological literature provides concrete ways to make outcome prediction robust to missing layers by treating absence as informative and by designing models that operate on whatever data are available.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app