Microsoft partnered with Surge AI to conduct blind human preference evaluations comparing MAI-Thinking-1 against Claude Sonnet 4.6 across over 1,200 tasks. Human raters preferred MAI-Thinking-1, showing how qualitative human judgment captures practical user experience better than automated benchmarks alone.