Multilingual and Informal Web Datasets for Robust Language Modeling

Large Language Models (LLMs) tend to lack textual, sociolinguistic, and cultural diversity due to being trained on homogeneous data sources. This study promotes the development of innovative datasets from marginal internet sites, such as anonymous imageboards (4chan, 2channel), niche web forums, YouTube comments, X posts, and Reddit sub-reddits, to overcome this shortfall. These datasets prioritize multilingual and multicultural features, reflecting spontaneous language use and distinctive sociocultural communication. Our solution involves web scraping and utilizes existing archives, with NLLB for multilingual translation support, and domain-specific BERT-based model training for improved text embeddings. We also describe building a custom language identification classifier and AI-generated text detector, both trained on a paired human-AI dataset augmented using the Qwen3 model, to enhance data quality and analytical potential for research in linguistic diversity and model robustness.

Paper

Full text

PDF

Multilingual and Informal Web Datasets for Robust Language Modeling

Semantic Scholar · Computer Science · 2025

Abstract

Large Language Models (LLMs) tend to lack textual, sociolinguistic, and cultural diversity due to being trained on homogeneous data sources. This study promotes the development of innovative datasets from marginal internet sites, such as anonymous imageboards (4chan, 2channel), niche web forums, YouTube comments, X posts, and Reddit sub-reddits, to overcome this shortfall. These datasets prioritize multilingual and multicultural features, reflecting spontaneous language use and distinctive sociocultural communication. Our solution involves web scraping and utilizes existing archives, with NLLB for multilingual translation support, and domain-specific BERT-based model training for improved text embeddings. We also describe building a custom language identification classifier and AI-generated text detector, both trained on a paired human-AI dataset augmented using the Qwen3 model, to enhance data quality and analytical potential for research in linguistic diversity and model robustness.

Similar papers

© 2026 NYSGPT2525 LLC