TY - JOUR
T1 - Statistical methods to harmonize electronic health record data across healthcare systems
T2 - case study and lessons learned
AU - Shi, Xu
AU - Zhai, Yuqi
AU - Yu, Xianshi
AU - Li, Xiaoou
AU - Hazlehurst, Brian L.
AU - Nyongesa, Denis B.
AU - Sapp, Daniel S.
AU - Williamson, Brian D.
AU - Carrell, David S.
AU - Healy, Luesa
AU - Cushing-Haugen, Kara L.
AU - Wong, Jenna
AU - Wang, Shirley V.
AU - Floyd, James S.
AU - Shattuck, Kathleen
AU - McGown, Samuel
AU - Alam, Sarah
AU - Hernández-Muñoz, José J.
AU - Li, Jie
AU - Ma, Yong
AU - Stojanovic, Danijela
AU - Raman, Sudha R.
AU - Davis, Sharon E.
AU - Cai, Tianxi
AU - Nelson, Jennifer C.
AU - Heagerty, Patrick J.
N1 - Publisher Copyright:
© The Author(s) 2026. Published by Oxford University Press.
PY - 2026/3
Y1 - 2026/3
N2 - Motivation: Although common data models for electronic health record (EHR) data can facilitate multi-site data organization and querying, the same medical event may still be coded differently between healthcare systems. In this paper, we present statistical methods to identify and mitigate coding discrepancies using summary-level data, and demonstrate these methods using data from two FDA Sentinel data partners: Kaiser Permanente Washington and Kaiser Permanente Northwest. Results: We first characterize differences in coding patterns, then compute a code mapping matrix to harmonize data between systems. Our findings reveal significant heterogeneity in coded EHR data, even after adopting a common data model with the same coding system, highlighting the importance of data harmonization before downstream analyses. Our study also demonstrates the effectiveness of the data harmonization approaches, which provide a foundational data quality step to promote semantic interoperability, enhance data integration, and improve the integrity of study conclusions. Availability and implementation: Computation prototypes, including R/Python codes and examples, are included in Section 7, available as supplementary data at Bioinformatics online and will be posted on GitHub upon publication.
AB - Motivation: Although common data models for electronic health record (EHR) data can facilitate multi-site data organization and querying, the same medical event may still be coded differently between healthcare systems. In this paper, we present statistical methods to identify and mitigate coding discrepancies using summary-level data, and demonstrate these methods using data from two FDA Sentinel data partners: Kaiser Permanente Washington and Kaiser Permanente Northwest. Results: We first characterize differences in coding patterns, then compute a code mapping matrix to harmonize data between systems. Our findings reveal significant heterogeneity in coded EHR data, even after adopting a common data model with the same coding system, highlighting the importance of data harmonization before downstream analyses. Our study also demonstrates the effectiveness of the data harmonization approaches, which provide a foundational data quality step to promote semantic interoperability, enhance data integration, and improve the integrity of study conclusions. Availability and implementation: Computation prototypes, including R/Python codes and examples, are included in Section 7, available as supplementary data at Bioinformatics online and will be posted on GitHub upon publication.
UR - https://www.scopus.com/pages/publications/105034220562
UR - https://www.scopus.com/pages/publications/105034220562#tab=citedBy
U2 - 10.1093/bioinformatics/btag107
DO - 10.1093/bioinformatics/btag107
M3 - Article
C2 - 41769828
AN - SCOPUS:105034220562
SN - 1367-4803
VL - 42
JO - Bioinformatics
JF - Bioinformatics
IS - 3
ER -