BACKGROUND Accurate differential diagnosis of complex neurological disorders remains challenging due to overlapping clinical features and heterogeneous disease presentations. Although large language models (LLMs) show promise in clinical reasoning, prior studies benchmark performance against clinician consensus rather than biological ground truth. A neuropathologically confirmed benchmark dataset for evaluating diagnostic AI in neurology is currently lacking.
Assessing AI and Neurologist Diagnostic Reasoning Against Neuropathological Ground Truth
calendar_today
July 10, 2026
person
Leng, Y., Noori, A., Dickson, J. R., Serrano-Pozo, A., Avetisyan, M., Rodriguez, D., Rosenberg, E. S., He, Y., Oakley, D. H., Khurana, V. S., Hyman, B. T., Frosch, M. P., Das, S.
domain
medrxiv