Digital Divides in Scene Recognition: Uncovering Socioeconomic Biases in Deep Learning Systems

Automatic scene classification has applications ranging from urban planning to autonomous driving, yet little is known about how well these systems work across social differences. We investigate explicit and implicit biases in deep learning architectures, including deep convolutional neural networks (dCNNs) and multimodal large language models (MLLMs). We examined nearly one million images from user-submitted photographs and Airbnb listings from over 200 countries as well as all 3320 US counties. To isolate scene-specific biases, we ensured no people were in any of the photos. We found significant explicit socioeconomic biases across all models, including lower classification accuracy, higher classification uncertainty, and increased tendencies to assign labels that could be offensive when applied to homes (e.g., “slum”) in images from homes with lower socioeconomic status. We also found significant implicit biases, with pictures from lower socioeconomic conditions more aligned with word embeddings from negative concepts. All trends were consistent across countries and within the diverse economic and racial landscapes of the United States. This research thus demonstrates a novel bias in computer vision, emphasizing the need for more inclusive and representative training datasets.

Paper

Similar papers

© 2026 NYSGPT2525 LLC