Privacy Leakage in Federated Learning in Radiology Reports: Comparative Evaluation of Tokenizer and Batch-Size Privacy Risks
At batch size 64 on the discharge dataset, accuracy was 64.7% (GPT-2), 70% (RadBERT), and 67.5% (LLaMA-2), decreasing to 27.3%, 28.5%, and 27.5% at batch size 256. The extent of such privacy leakage in FL applied to radiology reports, and the role of tokenizer design, remains unclear.