Alzheimer’s disease (AD) represents the most prevalent cause of dementia, accounting for an esti-
mated 60–70
%
of the 55 million dementia cases worldwide, with prevalence projected to exceed
130 million by 2050 as populations age [1, 2]. The clinical and economic burden is substantial—
global costs were estimated at
$
1.3 trillion in 2019—and recent therapeutic advances, including the
approval of anti-amyloid immunotherapies, have shifted the clinical imperative toward detection
in the earliest, most treatable stages of the disease. Despite this urgency, a substantial proportion
of patients remain undiagnosed until moderate dementia, when neuronal loss is already extensive
and disease-modifying interventions offer diminishing benefit [3].
Current diagnostic pathways rely predominantly on clinical assessment and cognitive screening,
with confirmatory testing based on amyloid positron emission tomography (PET) or cerebrospinal
fluid (CSF) measurements of amyloid-beta and phosphorylated tau. Although these biomarkers
achieve high diagnostic accuracy, they are constrained by high cost, limited availability, and inva-
siveness, which restricts their use in primary care and in low- and middle-income countries
where the greatest growth in dementia prevalence is expected [4, 5]. Structural magnetic reso-
nance imaging (MRI) is inexpensive, widely accessible, and non-invasive, yet the hippocampal at-
rophy and cortical thinning that characterise early AD are subtle and overlap considerably with
normal ageing, rendering conventional volumetry insufficient for reliable early classification.
Conversely, plasma proteomic assays are minimally invasive and increasingly reproducible, but
individually they capture only a partial view of the underlying pathology. These limitations sug-
gest that no single modality can fully capture the heterogeneous biological processes driving AD,
and argue for approaches that integrate complementary sources of information.
In this study, we introduce a multi-modal deep learning framework that fuses structural MRI and
plasma proteomic profiles using a cross-attention transformer architecture designed to model in-
ter-modal dependencies explicitly. We trained and validated the model on a multi-cohort dataset
comprising 4,213 participants from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) and
the Australian Imaging, Biomarkers and Lifestyle (AIBL) study, and evaluated generalisation on
two independent external cohorts. Our framework achieved a classification AUC of 0.91 (95
%
CI
0.89–0.93) for discriminating AD from cognitively normal controls, and—critically—identified par-
ticipants with mild cognitive impairment who converted to AD within three years with 78
%
sensi-
tivity at 82
%
specificity, a median of 18 months before clinical conversion. We further show that at-
tention-based fusion outperforms both unimodal baselines and simple concatenation strategies,
and we identify a set of plasma proteins whose interaction with hippocampal morphology contrib-
utes most to the model’s predictions. These findings demonstrate that integrated multi-modal
analysis can provide a scalable, accessible pathway to earlier and more accurate diagnosis of
Alzheimer’s disease.
1 / 1