Speech-to-text systems can have excellent Word Error Rate (WER) scores yet still disappoint users because perceived quality — determined by speaker accuracy, formatting, and entity recognition — matters more than raw transcription metrics. WER measures whether the right words appear in the transcript but ignores structural elements like speaker labels and punctuation. As WER converges across providers, perceived quality becomes the key competitive differentiator in voice AI products.