High Performance I/O For Large Scale Deep Learning

Training deep learning (DL) models on petascale datasets is essential for achieving competitive and state-of-the-art performance in applications such as speech, video analytics, and object recognition. However, existing distributed filesystems were not developed for the access patterns and usability requirements of DL jobs. In this paper, we describe AIStore, a highly scalable, easy-to-deploy storage system, and WebDataset, a standards-based storage format and library that permits efficient access to very large datasets. We compare system performance experimentally using image classification workloads and storing training data on a variety of backends, including local SSDs, single-node NFS, and two identical bare-metal clusters: HDFS and AIStore.

Paper

References (14)

08ImageNethttp://www.image-net.org
09A New Challenger: Western Digital Black 1TB NVMe M.2 SSD Reviewtechgage.com/article/western-digital-black-1tb-ssd-review/2
10Tarproc utilitiesgithub.com/tmbdev/tarproc
11Scaling Uber’s Apache Hadoop Distributed Filesystem for Growtheng.uber.com/scaling-hdfs
12libhdfs, a JNI based C API for Hadoop’s Distributed File Systemhadoop

Scroll for more · 2 remaining

Similar papers

© 2026 NYSGPT2525 LLC