Optimizing Deep Learning Recommender Systems' Training On CPU Cluster Architectures

During the last two years, the goal of many researchers has been to squeeze\nthe last bit of performance out of HPC system for AI tasks. Often this\ndiscussion is held in the context of how fast ResNet50 can be trained.\nUnfortunately, ResNet50 is no longer a representative workload in 2020. Thus,\nwe focus on Recommender Systems which account for most of the AI cycles in\ncloud computing centers. More specifically, we focus on Facebook's DLRM\nbenchmark. By enabling it to run on latest CPU hardware and software tailored\nfor HPC, we are able to achieve more than two-orders of magnitude improvement\nin performance (110x) on a single socket compared to the reference CPU\nimplementation, and high scaling efficiency up to 64 sockets, while fitting\nultra-large datasets. This paper discusses the optimization techniques for the\nvarious operators in DLRM and which component of the systems are stressed by\nthese different operators. The presented techniques are applicable to a broader\nset of DL workloads that pose the same scaling challenges/characteristics as\nDLRM.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC