Federated learning (FL) promises privacy-preserving model training across decentralized clients, but its empirical behavior under realistic conditions of statistical heterogeneity and formal privacy constraints remains under-characterized in controlled settings. This paper presents a systematic, simulation-based study that quantifies how decentralized training degrades along five axes — predictive accuracy, convergence stability, model calibration, communication cost, and the privacy–utility tradeoff — relative to a centralized baseline. Using Federated Averaging (FedAvg) over a neural classifier, we vary the degree of non-IID skew through a Dirichlet partitioning scheme, induce client imbalance, and inject Gaussian noise to emulate a differential-privacy mechanism. Across a 10-client testbed we observe that accuracy falls from 85.8% (centralized) to 81.9% under IID federation and to 71.0% under severe skew, while inter-client accuracy variance grows nearly three-fold and update oscillations intensify. Strong differential privacy (? = 5) reduces accuracy to 64.3% and sharply worsens calibration. We further compare candidate aggregation algorithms, justify FedAvg as the controlled reference, and introduce a composite Robustness Index (RI) that summarizes degradation jointly over heterogeneity and noise. The study moves beyond model-building toward an audit of the structural limits of decentralized machine learning.
Keywords : federated learning; data heterogeneity; differential privacy; model calibration; communication efficiency; FedAvg; non-IID data; trustworthy machine learning.
Author : Brinda S H Gangapatnam
Title : Federated Learning Under Data Heterogeneity: A Systematic Study of Privacy–Accuracy–Communication Tradeoffs
Volume/Issue : 2026;03(05)
Page No : 01-13