Quality of real-world data as a condition for the transferability of medical artificial intelligence models
https://doi.org/10.37489/2782-3784-myrwd-107
EDN: WPRZIX
Abstract
Relevance. The relevance of the topic is determined by the fact that digital solutions that have demonstrated high accuracy on curated research datasets are increasingly being proposed for use in healthcare organizations, where data are shaped by the clinical practices of a particular institution.
Objective. To analyze factors related to real-world data (RWD) that may lead to a decline in the performance of medical artificial intelligence (AI) models when they are transferred to new settings, and to substantiate approaches to their preliminary assessment before implementation in a healthcare organization.
Materials and methods. A scoping review was conducted to systematize publications on the applicability of medical AI models beyond the original dataset. The search was performed in Scopus, Web of Science Core Collection, PubMed, Google Scholar and eLIBRARY. RU for the period from 2016 to 2026; earlier sources were included when necessary.
Results. The review identified the main types of medical data, causes of model performance degradation and approaches to assessing model applicability in a new clinical setting. Particular attention was paid to data shift as a cause of reduced accuracy, impaired calibration and an increased number of clinically significant errors. Approaches to external and local validation, assessment of model robustness in subgroups and post-implementation monitoring were considered. Practical requirements were also proposed for describing model applicability in a scientific publication or validation report.
Conclusion. The ability of a medical AI system to maintain its effectiveness when used outside the original environment should be assessed as a separate criterion of its reliability. Validation reports should provide a detailed and reproducible description of data sources, outcome definitions, testing conditions and results obtained in a specific healthcare organization, as well as a justified strategy for subsequent monitoring and maintenance of system performance.
About the Authors
S. F. AbdulkerimovaRussian Federation
Selimat F. Abdulkerimova — Master's degree
Moscow
Competing Interests:
The authors declare that there is no conflict of interest in connection with this publication.
Yu. A. Danilova
Russian Federation
Yulia A. Danilova — Master's degree
Moscow
Competing Interests:
The authors declare that there is no conflict of interest in connection with this publication.
References
1. Gusev A.V., Zingerman B.V., Tyufilin D.S., Zinchenko V.V. Electronic medical records as a source of real-world clinical data. Real-World Data & Evidence. 2022;2(2):8-20. (In Russ.).
2. Borovskaya V.G., Gomon Y.M. Quality assessment of real-world data. Real-World Data & Evidence. 2022;2(4):10-16. (In Russ.).
3. Svechkareva I.R., Kurylev A.A., Shilova D.E. Particular qualities of evaluation of electronic card data in modern healthcare. Real-World Data & Evidence. 2024;4(4):28-34. (In Russ.).
4. Dmitrieva N.Y. Machine learning capabilities for the diagnosis of orphan diseases. Real-World Data & Evidence. 2023;3(3):36-39. (In Russ.).
5. Gribova V.V., Okun D.B., Shalfeeva E.A. Medical decision support model for diagnosing and rehabilitating patients with disabilities. Real-World Data & Evidence. 2025;5(4):81-96. (In Russ.).
6. Finlayson SG, Subbaswamy A, Singh K, et al. The Clinician and Dataset Shift in Artificial Intelligence. The New England Journal of Medicine. 2021 Jul;385(3):283-286. DOI: 10.1056/nejmc2104626.
7. Guan H, Guan H, Liu M. Domain Adaptation for Medical Image Analysis: A Survey. IEEE Transactions on Bio-medical Engineering. 2022 Mar;69(3):1173-1185. DOI: 10.1109/tbme.2021.3117407.
8. Zhou K, Liu Z, Qiao Y, et al. Domain Generalization: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2023 Apr;45(4):4396-4415. DOI: 10.1109/tpami.2022.3195549.
9. Zech JR, Badgeley MA, Liu M, et al. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. Plos Medicine. 2018 Nov;15(11):e1002683. DOI: 10.1371/journal.pmed.1002683.
10. DeGrave AJ, Janizek JD, Lee SI. AI for radiographic COVID-19 detection selects shortcuts over signal. medRxiv [Preprint]. 2020 Oct 7:2020.09.13.20193565. doi: 10.1101/2020. 09.13.20193565. Update in: This article has been published with doi: 10.1038/s42256-021-00338-7.
11. Ghafoorian M, Mehrtash A, Kapur T, et al. Transfer learning for domain adaptation in MRI: application in brain lesion segmentation. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2017. Cham: Springer; 2017:516–524. doi:10.1007/978-3-319-66179-7_59.
12. Daneshjou R, Vodrahalli K, Novoa RA, et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv. 2022 Aug 12;8(32):eabq6147. doi: 10.1126/sciadv.abq6147.
13. Seyyed-Kalantari L, Zhang H, McDermott MBA, et al. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021 Dec;27(12):2176-2182. doi: 10.1038/s41591-021-01595-0.
14. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019 Oct 25;366(6464):447-453. doi: 10.1126/science.aax2342.
15. Wong A, Otles E, Donnelly JP, et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Intern Med. 2021 Aug 1;181(8):1065-1070. doi: 10.1001/jamainternmed.2021.2626. Erratum in: JAMA Intern Med. 2021 Aug 1;181(8):1144. doi: 10.1001/jamainternmed.2021.3907.
16. Sendak MP, Ratliff W, Sarro D, et al. Real-World Integration of a Sepsis Deep Learning Technology Into Routine Clinical Care: Implementation Study. JMIR Med Inform. 2020 Jul 15;8(7):e15182. doi: 10.2196/15182.
17. Liu X, Glocker B, McCradden MM, et al. The medical algorithmic audit. Lancet Digit Health. 2022 May;4(5):e384-e397. doi: 10.1016/S2589-7500(22)00003-6. Epub 2022 Apr 5. Erratum in: Lancet Digit Health. 2022 Jun;4(6):e405. doi: 10.1016/S2589-7500(22)00089-9.
18. Tricco AC, Lillie E, Zarin W, et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med. 2018 Oct 2;169(7):467-473. doi: 10.7326/M18-0850.
19. Peters MDJ, Marnie C, Tricco AC, et al. Updated methodological guidance for the conduct of scoping reviews. JBI Evid Synth. 2020 Oct;18(10):2119-2126. doi: 10.11124/JBIES-20-00167.
20. Youssef A, Pencina M, Thakur A, et al. External validation of AI models in health should be replaced with recurring local validation. Nat Med. 2023 Nov;29(11):2686-2687. doi: 10.1038/s41591-023-02540-z.
21. Collins GS, Dhiman P, Ma J, et al. Evaluation of clinical prediction models (part 1): from development to external validation. BMJ. 2024 Jan 8;384:e074819. doi: 10.1136/bmj-2023-074819.
22. Riley RD, Archer L, Snell KIE, et al. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ. 2024 Jan 15;384:e074820. doi: 10.1136/bmj-2023-074820.
23. Van Calster B, McLernon DJ, van Smeden M, et al; Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019 Dec 16;17(1):230. doi: 10.1186/s12916-019-1466-7.
24. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006 Nov-Dec;26(6):565-74. doi: 10.1177/0272989X06295361.
25. Davis J, Goadrich M. The relationship between precision-recall and ROC curves. Proceedings of the 23rd International Conference on Machine Learning. 2006:233–240.
26. Koch LM, Baumgartner CF, Berens P. Distribution shift detection for the postmarket surveillance of medical AI algorithms: a retrospective simulation study. NPJ Digit Med. 2024 May 9;7(1):120. doi: 10.1038/s41746-024-01085-w.
27. Guo LL, Pfohl SR, Fries J, et al. Systematic Review of Approaches to Preserve Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine. Appl Clin Inform. 2021 Aug;12(4):808-815. doi: 10.1055/s-0041-1735184.
28. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024 Apr 16;385:e078378. doi: 10.1136/bmj-2023-078378. Erratum in: BMJ. 2024 Apr 18;385:q902. doi: 10.1136/bmj.q902.
29. Sounderajah V, Guni A, Liu X, et al; STARD-AI Steering Committee. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025 Oct;31(10):3283-3289. doi: 10.1038/s41591-025-03953-8. Epub 2025 Sep 15. Erratum in: Nat Med. 2026 Jul 13. doi: 10.1038/s41591-026-04570-9.
30. Liu X, Cruz Rivera S, Moher D, et al; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020 Sep;26(9):1364-1374. doi: 10.1038/s41591-020-1034-x.
31. Cruz Rivera S, Liu X, Chan AW, et al; SPIRIT-AI and CONSORT-AI Working Group; SPIRIT-AI and CONSORT-AI Steering Group; SPIRIT-AI and CONSORT-AI Consensus Group. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020 Sep;26(9):1351-1363. doi: 10.1038/s41591-020-1037-7.
32. Vasey B, Nagendran M, Campbell B, et al; DECIDE-AI expert group. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ. 2022 May 18;377:e070904. doi: 10.1136/bmj-2022-070904.
33. Mongan J, Moy L, Kahn CE Jr. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): A Guide for Authors and Reviewers. Radiol Artif Intell. 2020 Mar 25;2(2):e200029. doi: 10.1148/ryai.2020200029.
34. Norgeot B, Quer G, Beaulieu-Jones BK, et al. Minimum information about clinical artificial intelligence modeling: the MI-CLAIM checklist. Nat Med. 2020 Sep;26(9):1320-1324. doi: 10.1038/s41591-020-1041-y.
35. Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025 Mar 24;388:e082505. doi: 10.1136/bmj-2024-082505.
36. Wu E, Wu K, Daneshjou R, et al. How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals. Nat Med. 2021 Apr;27(4):582-584. doi: 10.1038/s41591-021-01312-x.
37. Grigoryev S.G., Lobzin Yu.V., Skripchenko N.V. The role and place of logistic regression and ROC analysis in solving medical diagnostic task. Journal Infectology. 2016;8(4):36-45. (In Russ.).
38. Abdulkerimova SF. Quality of Real-world Clinical Data as a Condition for Transferability of Medical Artificial Intelligence Models: Scoping Review Protocol and Search Record. OSF. 2026; doi:10.17605/OSF.IO/MDKUW.
Review
For citations:
Abdulkerimova S.F., Danilova Yu.A. Quality of real-world data as a condition for the transferability of medical artificial intelligence models. Real-World Data & Evidence. (In Russ.) https://doi.org/10.37489/2782-3784-myrwd-107. EDN: WPRZIX
JATS XML





















