A Computational Method for Learning Disease Trajectories from Partially Observable EHR Data

Wonsuk Oh, Michael S. Steinbach, M. Regina Castro, Kevin A. Peterson, Vipin Kumar, Pedro J. Caraballo, Gyorgy J. Simon

Research output: Contribution to journalArticlepeer-review

1 Scopus citations


Diseases can show different courses of progression even when patients share the same risk factors. Recent studies have revealed that the use of trajectories, the order in which diseases manifest throughout life, can be predictive of the course of progression. In this study, we propose a novel computational method for learning disease trajectories from EHR data. The proposed method consists of three parts: first, we propose an algorithm for extracting trajectories from EHR data; second, three criteria for filtering trajectories; and third, a likelihood function for assessing the risk of developing a set of outcomes given a trajectory set. We applied our methods to extract a set of disease trajectories from Mayo Clinic EHR data and evaluated it internally based on log-likelihood, which can be interpreted as the trajectories' ability to explain the observed (partial) disease progressions. We then externally evaluated the trajectories on EHR data from an independent health system, M Health Fairview. The proposed algorithm extracted a comprehensive set of disease trajectories that can explain the observed outcomes substantially better than competing methods and the proposed filtering criteria selected a small subset of disease trajectories that are highly interpretable and suffered only a minimal (relative 5%) loss of the ability to explain disease progression in both the internal and external validation.

Original languageEnglish (US)
Article number9456038
Pages (from-to)2476-2486
Number of pages11
JournalIEEE Journal of Biomedical and Health Informatics
Issue number7
StatePublished - Jul 1 2021

Bibliographical note

Funding Information:
Manuscript received August 24, 2020; revised November 16, 2020, December 28, 2020, April 19, 2021, and June 4, 2021; accepted June 6, 2021. Date of publication June 15, 2021; date of current version July 20, 2021. This work was supported by NIH Award LM011972, NSF Awards IIS 1602394 and IIS 1602198. The views expressed in this paper are those of the authors and do not necessarily reflect the views of the funding agencies. (Corresponding author: Gyorgy Simon.) Wonsuk Oh is with the Institute for Health Informatics, University of Minnesota, Minneapolis, MN 55455 USA, and also with Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, NY 10029 USA (e-mail: ohxxx215@umn.edu).

Publisher Copyright:
© 2013 IEEE.


  • Disease trajectories
  • electronic health records
  • machine learning

PubMed: MeSH publication types

  • Journal Article
  • Research Support, N.I.H., Extramural
  • Research Support, U.S. Gov't, Non-P.H.S.


Dive into the research topics of 'A Computational Method for Learning Disease Trajectories from Partially Observable EHR Data'. Together they form a unique fingerprint.

Cite this