Artificial intelligence (AI) has emerged as a transformative force in modern medicine, with
machine learning models increasingly capable of identifying complex patterns in clinical data
that elude traditional statistical approaches [1]. Among the most promising applications is
early disease detection, where AI-driven screening tools have demonstrated the ability to rec‐
ognize conditions such as diabetic retinopathy, cardiovascular disease, and several malignan‐
cies at stages when intervention is most effective [2, 3]. Despite these advances, the transla‐
tion of AI systems from controlled research environments into routine clinical practice re‐
mains limited, constrained by persistent concerns regarding generalizability, interpretability,
and regulatory approval [4]. Furthermore, existing models are frequently developed on homo‐
geneous patient populations, raising unresolved questions about their reliability across diverse
demographic groups and care settings [5]. This paper addresses these gaps by evaluating the
diagnostic accuracy of a novel deep learning framework designed for the early detection of
chronic kidney disease across multiple clinical cohorts. Drawing on electronic health records
from four geographically distinct healthcare systems, we assess whether the proposed model
maintains consistent performance across varying patient demographics, comorbid burdens,
and data quality levels. Our findings indicate that, although the model achieves state-of-the-
art overall accuracy, performance declines markedly among patients with multiple comorbidi‐
ties, underscoring the need for more granular and context-aware training strategies. These re‐
sults contribute to the growing body of evidence on the practical challenges of AI deployment
in clinical settings and provide a framework for developing more robust and equitable diag‐
nostic tools.
1 / 1