Jmir iconJmirSep 11, 2026 ~1 min source read

Quantifying the Impact of Anonymization-Induced Clinical Data Quality Loss: Methodological Quantitative Case Study Using Primary Diagnosis Codes and Hospital Length of Stay

The secondary use of electronic health record data requires robust privacy protection. A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce.

Share this story

Send the public story page.

Useful takeaways from this story.

The secondary use of electronic health record data requires robust privacy protection.

This study evaluated the analytical footprint of k-anonymity at k=5, 10, and 15 on 2 core data elements in retrospective hospital research: primary (ICD-10-GM) diagnosis codes, and hospital length of stay...

A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

The secondary use of electronic health record data requires robust privacy protection. A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce. This study evaluated the analytical footprint of k-anonymity at k=5, 10, and 15 on 2 core data elements in retrospective hospital research: primary (ICD-10-GM) diagnosis codes, and hospital length of stay (LOS).

How it works

  • It aimed to determine and quantify whether anonymization introduces meaningful distortions not captured by the anonymization tool itself, and whether diagnosis-specific LOS patterns remain reproducible...
  • Distributional distortion was assessed with the Kolmogorov-Smirnov statistic, quantile shifts, IQR changes, and tail changes.
  • Inferential reproducibility was assessed with a 3-level linear mixed model.
  • We compared the intraclass correlation coefficient and diagnosis-level effect concordance.
  • Concordance was quantified using Spearman ρ and Lin concordance correlation coefficient, both with 95% CIs.

Details worth keeping

Anonymization was performed with the ARX tool. It used record suppression and microaggregation. Categorical fidelity was assessed with the Jaccard coefficient and Cramer.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app