Understanding Depressive Symptoms and Psychosocial Stressors on Twitter: A Corpus-Based Study

Danielle Mowery; Hilary Smith; Tyler Cheney; Greg Stoddard; Glen Coppersmith; Craig Bryan; Mike Conway

doi:10.2196/jmir.6895

Understanding Depressive Symptoms and Psychosocial Stressors on Twitter: A Corpus-Based Study

J Med Internet Res. 2017 Feb 28;19(2):e48. doi: 10.2196/jmir.6895.

Authors

Danielle Mowery¹, Hilary Smith², Tyler Cheney², Greg Stoddard³, Glen Coppersmith^{4

5}, Craig Bryan², Mike Conway¹

Affiliations

¹ Department of Biomedical Informatics, University of Utah, Salt Lake City, UT, United States.
² Department of Psychology, University of Utah, Salt Lake City, UT, United States.
³ Department of Family And Preventive Medicine, University of Utah, Salt Lake City, UT, United States.
⁴ Qntfy, Crownsville, MD, United States.
⁵ Human Language Technology Center of Excellence, John Hopkins University, Baltimore, MD, United States.

Abstract

Background: With a lifetime prevalence of 16.2%, major depressive disorder is the fifth biggest contributor to the disease burden in the United States.

Objective: The aim of this study, building on previous work qualitatively analyzing depression-related Twitter data, was to describe the development of a comprehensive annotation scheme (ie, coding scheme) for manually annotating Twitter data with Diagnostic and Statistical Manual of Mental Disorders, Edition 5 (DSM 5) major depressive symptoms (eg, depressed mood, weight change, psychomotor agitation, or retardation) and Diagnostic and Statistical Manual of Mental Disorders, Edition IV (DSM-IV) psychosocial stressors (eg, educational problems, problems with primary support group, housing problems).

Methods: Using this annotation scheme, we developed an annotated corpus, Depressive Symptom and Psychosocial Stressors Acquired Depression, the SAD corpus, consisting of 9300 tweets randomly sampled from the Twitter application programming interface (API) using depression-related keywords (eg, depressed, gloomy, grief). An analysis of our annotated corpus yielded several key results.

Results: First, 72.09% (6829/9473) of tweets containing relevant keywords were nonindicative of depressive symptoms (eg, "we're in for a new economic depression"). Second, the most prevalent symptoms in our dataset were depressed mood and fatigue or loss of energy. Third, less than 2% of tweets contained more than one depression related category (eg, diminished ability to think or concentrate, depressed mood). Finally, we found very high positive correlations between some depression-related symptoms in our annotated dataset (eg, fatigue or loss of energy and educational problems; educational problems and diminished ability to think).

Conclusions: We successfully developed an annotation scheme and an annotated corpus, the SAD corpus, consisting of 9300 tweets randomly-selected from the Twitter application programming interface using depression-related keywords. Our analyses suggest that keyword queries alone might not be suitable for public health monitoring because context can change the meaning of keyword in a statement. However, postprocessing approaches could be useful for reducing the noise and improving the signal needed to detect depression symptoms using social media.

Keywords: Twitter messaging; data annotation; machine learning; major depressive disorder; natural language processing; social media.

©Danielle Mowery, Hilary Smith, Tyler Cheney, Greg Stoddard, Glen Coppersmith, Craig Bryan, Mike Conway. Originally published in the Journal of Medical Internet Research (http://www.jmir.org), 28.02.2017.

MeSH terms

Depression / diagnosis*
Depression / epidemiology
Depressive Disorder, Major / diagnosis*
Depressive Disorder, Major / epidemiology
Humans
Internet / statistics & numerical data*
Machine Learning
Psychology
Social Media / statistics & numerical data*
Stress, Psychological / diagnosis*
Stress, Psychological / epidemiology

Abstract

MeSH terms

Grants and funding