Using Robust PCA to estimate regional characteristics of language use from geo-tagged Twitter messages

Principal component analysis (PCA) and related techniques have been\nsuccessfully employed in natural language processing. Text mining applications\nin the age of the online social media (OSM) face new challenges due to\nproperties specific to these use cases (e.g. spelling issues specific to texts\nposted by users, the presence of spammers and bots, service announcements,\netc.). In this paper, we employ a Robust PCA technique to separate typical\noutliers and highly localized topics from the low-dimensional structure present\nin language use in online social networks. Our focus is on identifying\ngeospatial features among the messages posted by the users of the Twitter\nmicroblogging service. Using a dataset which consists of over 200 million\ngeolocated tweets collected over the course of a year, we investigate whether\nthe information present in word usage frequencies can be used to identify\nregional features of language use and topics of interest. Using the PCA pursuit\nmethod, we are able to identify important low-dimensional features, which\nconstitute smoothly varying functions of the geographic location.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC