The goal from this study is to identify suicide risk factors on Twitter. We propose a machine learning framework that could be potentially useful for suicide prevention interventions. We applied search terms from the suicidal ideation tracking framework pro-posed by Jashinsky et al. and downloaded 12,066 public tweets from 3,873 users via Twitter's application programming interface (API). We created "HighRisk" or "AtRisk" labels for users based on their suicidal ideation terms' usage and applied three topic discovery algorithms to find underlying suicide risk factors among users, which were subsequently used to classify users into "HighRisk" or "AtRisk". Algorithms applied included Latent Semantic Analysis, Latent Dirichlet Allocation, Non-negative Matrix Factorization, Decision Tree and K-means Clustering. Our topic discovery approach detected 7 out of 12 suicide risk factors proposed by Jashinsky et al. Using a decision tree classification model that utilized these factors, we achieved 0.844 in precision, 0.912 in sensitivity, and 0.829 in specificity in classifying users into "HighRisk" and "AtRisk" groups. The development of this framework supplements suicide researchers and suicide preven-tion efforts, with a potential to be employed at run-time.
Paper
Full text
Using Machine Learning Algorithms to Detect Suicide Risk Factors on Twitter
Semantic Scholar · Computer Science · 2019
Abstract
The goal from this study is to identify suicide risk factors on Twitter. We propose a machine learning framework that could be potentially useful for suicide prevention interventions. We applied search terms from the suicidal ideation tracking framework pro-posed by Jashinsky et al. and downloaded 12,066 public tweets from 3,873 users via Twitter's application programming interface (API). We created "HighRisk" or "AtRisk" labels for users based on their suicidal ideation terms' usage and applied three topic discovery algorithms to find underlying suicide risk factors among users, which were subsequently used to classify users into "HighRisk" or "AtRisk". Algorithms applied included Latent Semantic Analysis, Latent Dirichlet Allocation, Non-negative Matrix Factorization, Decision Tree and K-means Clustering. Our topic discovery approach detected 7 out of 12 suicide risk factors proposed by Jashinsky et al. Using a decision tree classification model that utilized these factors, we achieved 0.844 in precision, 0.912 in sensitivity, and 0.829 in specificity in classifying users into "HighRisk" and "AtRisk" groups. The development of this framework supplements suicide researchers and suicide preven-tion efforts, with a potential to be employed at run-time.