Shopify’s Sidekick AI assistant lacked training examples for refusing impossible requests because production logs only captured successful queries. The team built an automated curation pipeline using four frontier LLMs as judges, calibrated on 602 refusal annotations and requiring unanimous consensus, reaching 86.3% refusal accuracy with a 4.6% false positive rate. The approach lifted the customer segmentation skill’s evaluation score from 0.619 to 0.798, a 28.9% relative gain.