Distributed or federated training where some workers send bad gradients, by mistake or on purpose. This is similar to resilient consensus: the server, or each peer, has to aggregate vectors it can’t fully trust.

Aggregators

  • Krum (Blanchard, El Mhamdi, Guerraoui and Stainer, 2017) picks the gradient closest to its nearest neighbours.
  • Coordinate-wise median and trimmed mean (Yin, Chen, Ramchandran and Bartlett, 2018), with statistical error rates. The trimmed mean is a close cousin of W-MSR.

Relevance

Robot teams that learn together (shared maps, shared policies) have the same problem, with less compute, less reliable links, and physical consequences for mistakes.

References

  • P. Blanchard, E. M. El Mhamdi, R. Guerraoui, J. Stainer. Machine learning with adversaries: Byzantine tolerant gradient descent. NeurIPS, 2017.
  • D. Yin, Y. Chen, K. Ramchandran, P. Bartlett. Byzantine-robust distributed learning: towards optimal statistical rates. ICML, 2018.