Distributed or federated training where some workers send bad gradients, by mistake or on purpose. This is similar to resilient consensus: the server, or each peer, has to aggregate vectors it can’t fully trust.
Aggregators
- Krum (Blanchard, El Mhamdi, Guerraoui and Stainer, 2017) picks the gradient closest to its nearest neighbours.
- Coordinate-wise median and trimmed mean (Yin, Chen, Ramchandran and Bartlett, 2018), with statistical error rates. The trimmed mean is a close cousin of W-MSR.
Relevance
Robot teams that learn together (shared maps, shared policies) have the same problem, with less compute, less reliable links, and physical consequences for mistakes.
References
- P. Blanchard, E. M. El Mhamdi, R. Guerraoui, J. Stainer. Machine learning with adversaries: Byzantine tolerant gradient descent. NeurIPS, 2017.
- D. Yin, Y. Chen, K. Ramchandran, P. Bartlett. Byzantine-robust distributed learning: towards optimal statistical rates. ICML, 2018.