The piece outlines four hierarchical levels for testing legal AI systems, ranked by credibility for buyers. It explains how each testing approach, from internal benchmarking to real work product comparison, has distinct strengths and limitations for evaluating whether AI tools perform effectively in actual legal practice.