What's Been Happening in the Romanian News Landscape? A Detailed Analysis Grounded in Natural Language Processing Techniques
People strive to be connected to events happening worldwide in terms of politics, technology, sports, business, and many other domains. The main source of news today resides in online publications which can strongly influence the public opinion. Our purpose is to build a comprehensive automated pipeline, integrating various Natural Language Processing techniques, to process online news written in the Romanian language. Our dataset consists of 631,565 news articles from various Romanian publications between May 2004 and December 2019 which are used to detect semantic similarities between articles and rank various publications in terms of their influence. Furthermore, we created visualizations to ease the understanding of results and ensure efficient text retrieval over the gathered articles. In the future, we plan to apply opinion mining, geographical names extraction and content quality assessments relating, for example, to the likelihood of being a fake news.
Paper
Full text
What's Been Happening in the Romanian News Landscape? A Detailed Analysis Grounded in Natural Language Processing Techniques
OpenAlex · Advanced Text Analysis Techniques · 2020
Abstract
People strive to be connected to events happening worldwide in terms of politics, technology, sports, business, and many other domains. The main source of news today resides in online publications which can strongly influence the public opinion. Our purpose is to build a comprehensive automated pipeline, integrating various Natural Language Processing techniques, to process online news written in the Romanian language. Our dataset consists of 631,565 news articles from various Romanian publications between May 2004 and December 2019 which are used to detect semantic similarities between articles and rank various publications in terms of their influence. Furthermore, we created visualizations to ease the understanding of results and ensure efficient text retrieval over the gathered articles. In the future, we plan to apply opinion mining, geographical names extraction and content quality assessments relating, for example, to the likelihood of being a fake news.