Using Machine Learning Algorithms to Detect Suicide Risk Factors on Twitter

The goal from this study is to identify suicide risk factors on Twitter. We propose a machine learning framework that could be potentially useful for suicide prevention interventions. We applied search terms from the suicidal ideation tracking framework pro-posed by Jashinsky et al. and downloaded 12,066 public tweets from 3,873 users via Twitter's application programming interface (API). We created "HighRisk" or "AtRisk" labels for users based on their suicidal ideation terms' usage and applied three topic discovery algorithms to find underlying suicide risk factors among users, which were subsequently used to classify users into "HighRisk" or "AtRisk". Algorithms applied included Latent Semantic Analysis, Latent Dirichlet Allocation, Non-negative Matrix Factorization, Decision Tree and K-means Clustering. Our topic discovery approach detected 7 out of 12 suicide risk factors proposed by Jashinsky et al. Using a decision tree classification model that utilized these factors, we achieved 0.844 in precision, 0.912 in sensitivity, and 0.829 in specificity in classifying users into "HighRisk" and "AtRisk" groups. The development of this framework supplements suicide researchers and suicide preven-tion efforts, with a potential to be employed at run-time.

Paper

Full text

PDF

Using Machine Learning Algorithms to Detect Suicide Risk Factors on Twitter

Semantic Scholar · Computer Science · 2019

Abstract

The goal from this study is to identify suicide risk factors on Twitter. We propose a machine learning framework that could be potentially useful for suicide prevention interventions. We applied search terms from the suicidal ideation tracking framework pro-posed by Jashinsky et al. and downloaded 12,066 public tweets from 3,873 users via Twitter's application programming interface (API). We created "HighRisk" or "AtRisk" labels for users based on their suicidal ideation terms' usage and applied three topic discovery algorithms to find underlying suicide risk factors among users, which were subsequently used to classify users into "HighRisk" or "AtRisk". Algorithms applied included Latent Semantic Analysis, Latent Dirichlet Allocation, Non-negative Matrix Factorization, Decision Tree and K-means Clustering. Our topic discovery approach detected 7 out of 12 suicide risk factors proposed by Jashinsky et al. Using a decision tree classification model that utilized these factors, we achieved 0.844 in precision, 0.912 in sensitivity, and 0.829 in specificity in classifying users into "HighRisk" and "AtRisk" groups. The development of this framework supplements suicide researchers and suicide preven-tion efforts, with a potential to be employed at run-time.

Similar papers

© 2026 NYSGPT2525 LLC