Clinical AI Trust Half-Life: Time-Decaying Prediction Assurance Under Data and Fairness Drift

Authors

  • Nithesh Naik Jesselton University College, Kota Kinabalu, Sabah Author
  • Desmond Cherng En Lee Jesselton University College, Kota Kinabalu, Sabah Author
  • Harikrishnan A/L Ramiah Department of Electrical Engineering, Faculty of Engineering, Universiti Malaya, Kuala Lumpur Author

Keywords:

Clinical Artificial Intelligence, Data Drift, Model Calibration, Fairness Drift, Post-Deployment Monitoring, Prediction Assurance, Trust Half-Life

Abstract

Clinical artificial intelligence (AI) models are commonly validated at a fixed point, although the conditions that support their predictions can change after deployment. This paper introduces Clinical AI Trust Half-Life, a theoretical measure of the time required for overall prediction assurance to decline to one-half of its initially validated level. Assurance is represented as a normalized composite of predictive performance, calibration, distribution stability, and subgroup fairness, while workflow change and monitoring delay influence the rate and practical consequences of decay. A reproducible simulation trains logistic regression on synthetic baseline data for two patient groups and then evaluates the unchanged model across thirteen periods with gradual covariate drift, increasing outcome prevalence, changing predictor–outcome relationships, and faster degradation in one subgroup. In the executed simulation, assurance decreased from 0.858 at baseline to 0.334 at period 12. The empirical half-life occurred at period 9, when assurance first fell below 0.429; an exponential approximation yielded 8.78 periods. The simulation demonstrates that aggregate discrimination can remain superficially usable while calibration and subgroup sensitivity deteriorate substantially. Trust half-life is proposed as a governance summary, not a substitute for component-level surveillance or clinical validation. It can support risk-adjusted monitoring intervals, early-warning thresholds, subgroup audits, revalidation, recalibration, and explicit assignment of maintenance responsibilities.

References

Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019;17:195. doi:10.1186/s12916-019-1426-2.

Subbaswamy A, Saria S. From development to deployment: dataset shift, causality, and shift-stable models in health AI. Biostatistics. 2020;21(2):345-352. doi:10.1093/biostatistics/kxz041.

Nestor B, McDermott MBA, Boag W, Berner G, Naumann T, Hughes MC, et al. Feature robustness in non-stationary health records: caveats to deployable model performance in common clinical machine learning tasks. Proc Mach Learn Res. 2019;106:381-405.

Guo LL, Pfohl SR, Fries J, Posada J, Fleming SL, Aftandilian C, et al. Systematic review of approaches to preserve machine learning performance in the presence of temporal dataset shift in clinical medicine. Appl Clin Inform. 2021;12(4):808-815. doi:10.1055/s-0041-1735184.

Guo LL, Pfohl SR, Fries J, Johnson AEW, Posada J, Aftandilian C, et al. Evaluation of domain generalization and adaptation on improving model robustness to temporal dataset shift in clinical medicine. Sci Rep. 2022;12:2726. doi:10.1038/s41598-022-06484-1.

Sahiner B, Chen W, Samala RK, Petrick N. Data drift in medical machine learning: implications and potential remedies. Br J Radiol. 2023;96(1150):20220878. doi:10.1259/bjr.20220878.

Davis SE, Walsh CG, Matheny ME. Open questions and research gaps for monitoring and updating AI-enabled tools in clinical settings. Front Digit Health. 2022;4:958284. doi:10.3389/fdgth.2022.958284.

Kore A, Abbasi Bavil E, Subasri V, Abdalla M, Fine B, Dolatabadi E, et al. Empirical data drift detection experiments on real-world medical imaging data. Nat Commun. 2024;15:1887. doi:10.1038/s41467-024-46142-w.

Koch LM, Baumgartner CF, Berens P. Distribution shift detection for the postmarket surveillance of medical AI algorithms: a retrospective simulation study. NPJ Digit Med. 2024;7:120. doi:10.1038/s41746-024-01085-w.

Subasri V, Krishnan A, Kore A, Dhalla A, Pandya D, Wang B, et al. Detecting and remediating harmful data shifts for the responsible deployment of clinical AI models. JAMA Netw Open. 2025;8(6):e2513685. doi:10.1001/jamanetworkopen.2025.13685.

Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW; STRATOS Topic Group. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17:230. doi:10.1186/s12916-019-1466-7.

Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330.

Rajkomar A, Hardt M, Howell MD, Corrado G, Chin MH. Ensuring fairness in machine learning to advance health equity. Ann Intern Med. 2018;169(12):866-872. doi:10.7326/M18-1990.

Seyyed-Kalantari L, Zhang H, McDermott MBA, Chen IY, Ghassemi M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med. 2021;27(12):2176-2182. doi:10.1038/s41591-021-01595-0.

Chen RJ, Wang JJ, Williamson DFK, Chen TY, Lipkova J, Lu MY, et al. Algorithm fairness in artificial intelligence for medicine and healthcare. Nat Biomed Eng. 2023;7(6):719-742. doi:10.1038/s41551-023-01056-8.

Davis SE, Dorn C, Park DJ, Matheny ME. Emerging algorithmic bias: fairness drift as the next dimension of model maintenance and sustainability. J Am Med Inform Assoc. 2025;32(5):845-854. doi:10.1093/jamia/ocaf039.

Ramspek CL, Jager KJ, Dekker FW, Zoccali C, van Diepen M. External validation of prognostic models: what, why, how, when and where? Clin Kidney J. 2021;14(1):49-58. doi:10.1093/ckj/sfaa188.

Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, et al.; DECIDE-AI expert group. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ. 2022;377:e070904. doi:10.1136/bmj-2022-070904.

Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Lancet Digit Health. 2020;2(10):e537-e548. doi:10.1016/S2589-7500(20)30218-1.

Kim GYE, Corbin CK, Grolleau F, Baiocchi M, Chen JH. Monitoring strategies for continuous evaluation of deployed clinical prediction models. J Biomed Inform. 2025;168:104854. doi:10.1016/j.jbi.2025.104854.

Pfohl SR, Zhang H, Xu Y, Foryciarz A, Ghassemi M, Shah NH. A comparison of approaches to improve worst-case predictive model performance over patient subpopulations. Sci Rep. 2022;12:3254. doi:10.1038/s41598-022-07167-7.

Celi LA, Cellini J, Charpignon ML, Dee EC, Dernoncourt F, Eber R, et al.; MIT Critical Data. Sources of bias in artificial intelligence that perpetuate healthcare disparities-a global review. PLOS Digit Health. 2022;1(3):e0000022. doi:10.1371/journal.pdig.0000022.

Duckworth C, Chmiel FP, Burns DK, Zlatev ZD, White NM, Daniels TWV, et al. Using explainable machine learning to characterise data drift and detect emergent health risks for emergency department admissions during COVID-19. Sci Rep. 2021;11:23017. doi:10.1038/s41598-021-02481-y.

Downloads

Published

2026-09-30

Issue

Section

Articles

How to Cite

Clinical AI Trust Half-Life: Time-Decaying Prediction Assurance Under Data and Fairness Drift. (2026). Journal of Healthcare Analytics and Informatics, 1(1), 1-9. https://landing.wrunion.org/ojs/index.php/jhai/article/view/28