Data-driven automated classification algorithms for acute health conditions: applying PheNorm to COVID-19 disease

Joshua C Smith; Brian D Williamson; David J Cronkite; Daniel Park; Jill M Whitaker; Michael F McLemore; Joshua T Osmanski; Robert Winter; Arvind Ramaprasan; Ann Kelley; Mary Shea; Saranrat Wittayanukorn; Danijela Stojanovic; Yueqin Zhao; Sengwee Toh; Kevin B Johnson; David M Aronoff; David S Carrell

doi:10.1093/jamia/ocad241

Data-driven automated classification algorithms for acute health conditions: applying PheNorm to COVID-19 disease

J Am Med Inform Assoc. 2024 Feb 16;31(3):574-582. doi: 10.1093/jamia/ocad241.

Authors

Joshua C Smith¹, Brian D Williamson², David J Cronkite², Daniel Park¹, Jill M Whitaker¹, Michael F McLemore¹, Joshua T Osmanski¹, Robert Winter¹, Arvind Ramaprasan², Ann Kelley², Mary Shea², Saranrat Wittayanukorn³, Danijela Stojanovic³, Yueqin Zhao³, Sengwee Toh⁴, Kevin B Johnson⁵, David M Aronoff⁶, David S Carrell²

Affiliations

¹ Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37203, United States.
² Kaiser Permanente Washington Health Research Institute, Seattle, WA 98101, United States.
³ Center for Drug Evaluation and Research, US Food and Drug Administration, Silver Spring, MD 20903, United States.
⁴ Harvard Pilgrim Health Care Institute, Boston, MA 02215, United States.
⁵ Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA 19104, United States.
⁶ Department of Medicine, Indiana University School of Medicine, Indianapolis, IN 46202, United States.

PMID: 38109888
PMCID: PMC10873852 (available on 2024-12-18)
DOI: 10.1093/jamia/ocad241

Abstract

Objectives: Automated phenotyping algorithms can reduce development time and operator dependence compared to manually developed algorithms. One such approach, PheNorm, has performed well for identifying chronic health conditions, but its performance for acute conditions is largely unknown. Herein, we implement and evaluate PheNorm applied to symptomatic COVID-19 disease to investigate its potential feasibility for rapid phenotyping of acute health conditions.

Materials and methods: PheNorm is a general-purpose automated approach to creating computable phenotype algorithms based on natural language processing, machine learning, and (low cost) silver-standard training labels. We applied PheNorm to cohorts of potential COVID-19 patients from 2 institutions and used gold-standard manual chart review data to investigate the impact on performance of alternative feature engineering options and implementing externally trained models without local retraining.

Results: Models at each institution achieved AUC, sensitivity, and positive predictive value of 0.853, 0.879, 0.851 and 0.804, 0.976, and 0.885, respectively, at quantiles of model-predicted risk that maximize F1. We report performance metrics for all combinations of silver labels, feature engineering options, and models trained internally versus externally.

Discussion: Phenotyping algorithms developed using PheNorm performed well at both institutions. Performance varied with different silver-standard labels and feature engineering options. Models developed locally at one site also worked well when implemented externally at the other site.

Conclusion: PheNorm models successfully identified an acute health condition, symptomatic COVID-19. The simplicity of the PheNorm approach allows it to be applied at multiple study sites with substantially reduced overhead compared to traditional approaches.

Keywords: COVID-19; electronic health records; machine learning; natural language processing; phenotyping.

MeSH terms

Algorithms*
COVID-19*
Electronic Health Records
Humans
Machine Learning
Natural Language Processing

Abstract

MeSH terms

Grants and funding