Jmir iconJmirSep 11, 2026 ~1 min source read

Privacy Leakage in Federated Learning in Radiology Reports: Comparative Evaluation of Tokenizer and Batch-Size Privacy Risks

At batch size 64 on the discharge dataset, accuracy was 64.7% (GPT-2), 70% (RadBERT), and 67.5% (LLaMA-2), decreasing to 27.3%, 28.5%, and 27.5% at batch size 256. The extent of such privacy leakage in FL applied to radiology reports, and the role of tokenizer design, remains unclear.

Share this story

Send the public story page.

Useful takeaways from this story.

The extent of such privacy leakage in FL applied to radiology reports, and the role of tokenizer design, remains unclear.

Models were trained using 3 tokenizers (GPT-2, RadBERT, and LLaMA-2) with batch sizes of 64, 128, and 256.

At batch size 64 on the discharge dataset, accuracy was 64.7% (GPT-2), 70% (RadBERT), and 67.5% (LLaMA-2), decreasing to 27.3%, 28.5%, and 27.5% at batch size 256.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

The extent of such privacy leakage in FL applied to radiology reports, and the role of tokenizer design, remains unclear. This study aimed to quantify gradient-based reconstruction of radiology report text in an FL setting and to compare privacy risk across 3 transformer tokenization strategies in a controlled, tokenizer-aware evaluation. Models were trained using 3 tokenizers (GPT-2, RadBERT, and LLaMA-2) with batch sizes of 64, 128, and 256.

How it works

  • At batch size 64 on the discharge dataset, accuracy was 64.7% (GPT-2), 70% (RadBERT), and 67.5% (LLaMA-2), decreasing to 27.3%, 28.5%, and 27.5% at batch size 256.
  • An active malicious-server threat model was assumed, and analytic gradient inversion was applied to recover text.
  • Reconstruction fidelity was measured over 5 runs using exact sentence accuracy, sentence-level bilingual evaluation understudy (S-BLEU), and recall-oriented understudy for gisting evaluation (ROUGE-L).

Details worth keeping

Six FL clients trained a GPT-2–style transformer (sequence length 32) on 2 public clinical-text corpora comprising 368,751 diagnostic reports, 98,206 discharge summaries, and 1500 MIMIC-CXR (Medical Information Mart for Intensive Care Chest X-Ray) radiology reports. Batch size was the dominant factor governing leakage.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app