Challenges associated with missing data in electronic health records: A case study of a risk prediction model for diabetes using data from Slovenian primary care.

Stiglic G.; Kocbek P.; Fijacko N.; Sheikh A.; Pajnkihar M.

Challenges associated with missing data in electronic health records: A case study of a risk prediction model for diabetes using data from Slovenian primary care.

Stiglic G., Kocbek P., Fijacko N., Sheikh A., Pajnkihar M.

The increasing availability of data stored in electronic health records brings substantial opportunities for advancing patient care and population health. This is, however, fundamentally dependant on the completeness and quality of data in these electronic health records. We sought to use electronic health record data to populate a risk prediction model for identifying patients with undiagnosed type 2 diabetes mellitus. We, however, found substantial (up to 90%) amounts of missing data in some healthcare centres. Attempts at imputing for these missing data or using reduced dataset by removing incomplete records resulted in a major deterioration in the performance of the prediction model. This case study illustrates the substantial wasted opportunities resulting from incomplete records by simulation of missing and incomplete records in predictive modelling process. Government and professional bodies need to prioritise efforts to address these data shortcomings in order to ensure that electronic health record data are maximally exploited for patient and population benefit.

Original publication

DOI

10.1177/1460458217733288

Type

Journal article

Journal

Health informatics journal

Publication Date

09/2019

Volume

Pages

951 - 959

Keywords

Humans, Diabetes Mellitus, Type 2, Risk Assessment, Case-Control Studies, Cross-Sectional Studies, Middle Aged, Primary Health Care, Slovenia, Female, Male, Electronic Health Records, Quality Improvement, Surveys and Questionnaires, Data Accuracy

Cookies on this website

Challenges associated with missing data in electronic health records: A case study of a risk prediction model for diabetes using data from Slovenian primary care.

Stiglic G., Kocbek P., Fijacko N., Sheikh A., Pajnkihar M.

DOI

Type

Journal

Publication Date

Volume

Pages

Keywords