Federated learning with hierarchical clustering of local updates to improve training on non-IID data
Federated learning (FL) is a well established method for performing machine\nlearning tasks over massively distributed data. However in settings where data\nis distributed in a non-iid (not independent and identically distributed)\nfashion -- as is typical in real world situations -- the joint model produced\nby FL suffers in terms of test set accuracy and/or communication costs compared\nto training on iid data. We show that learning a single joint model is often\nnot optimal in the presence of certain types of non-iid data. In this work we\npresent a modification to FL by introducing a hierarchical clustering step\n(FL+HC) to separate clusters of clients by the similarity of their local\nupdates to the global joint model. Once separated, the clusters are trained\nindependently and in parallel on specialised models. We present a robust\nempirical analysis of the hyperparameters for FL+HC for several iid and non-iid\nsettings. We show how FL+HC allows model training to converge in fewer\ncommunication rounds (significantly so under some non-iid settings) compared to\nFL without clustering. Additionally, FL+HC allows for a greater percentage of\nclients to reach a target accuracy compared to standard FL. Finally we make\nsuggestions for good default hyperparameters to promote superior performing\nspecialised models without modifying the the underlying federated learning\ncommunication protocol.\n
Paper
References (22)
Scroll for more · 10 remaining