General Symptom Extraction from VA Electronic Medical Notes

Guy Divita; Gang Luo; Le-Thuy T Tran; T Elizabeth Workman; Adi V Gundlapalli; Matthew H Samore

General Symptom Extraction from VA Electronic Medical Notes

Stud Health Technol Inform. 2017:245:356-360.

Authors

Guy Divita¹, Gang Luo², Le-Thuy T Tran¹, T Elizabeth Workman¹, Adi V Gundlapalli¹, Matthew H Samore¹

Affiliations

¹ VA Salt Lake City Health Care System, Salt Lake City, Utah, USA.
² Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, USA.

PMID: 29295115

Abstract

There is need for cataloging signs and symptoms, but not all are documented in structured data. The text from clinical records are an additional source of signs and symptoms. We describe a Natural Language Processing (NLP) technique to identify symptoms from text. Using a human-annotated reference corpus from VA electronic medical notes we trained and tested an NLP pipeline to identify and categorize symptoms. The technique includes a model created from an automatic machine learning model selection tool. Tested on a hold-out set, its precision at the mention level was 0.80, recall 0.74 and an overall f-score of 0.80. The tool was scaled-up to process a large corpus of 964,105 patient records.

Keywords: Diagnosis; Machine Learning; Natural Language Processing.

MeSH terms

Data Mining*
Electronic Health Records
Humans
Machine Learning*
Natural Language Processing*