Total: 1
The ADReSS and ADReSSo challenge datasets have become the de-facto standards for research on dementia detection through speech, with over half of recent ICASSP and Interspeech studies on the topic relying on them. Despite their widespread adoption, both datasets exhibit properties that may undermine the validity of reported results. In this paper, the acoustic variability in the test sets of ADReSS and ADReSSo is systematically examined. Results show that (i) near state-of-the-art classification performance on the test sets can be achieved using only two low-level acoustic features, (ii) classifiers trained on randomly permuted dementia labels can nonetheless reach competitive test performance, and (iii) feature discriminability does not generalize across Monte Carlo resampling, with no stable features emerging from low-level acoustic feature sets. These findings indicate that strong results on these datasets can stem from spurious correlations rather than pathology-relevant cues. The analysis calls for more cautious interpretation of benchmark performance, as well as renewed attention to small dataset evaluation strategies in dementia detection.