I present a new way to parallelize the training of convolutional neural networks across multiple GPUs. The method scales significantly better than all alternatives when applied to modern convolutional neural networks.
Paper
References (8)
08One of the workers sends its last-stage convolutional layer activities to all other workersparallel with this computation