As AI agents become a key persona for interacting with business services, the efficiency of Model Context Protocol (MCP) servers is paramount. This article, based on a talk by Fred from ALP, explores how to design and operate MCP servers to maximize “task completion rate.” It delves into the challenges of measuring this metric in production and proposes using proxy metrics like token cost and latency. The article then provides actionable insights and patterns for MCP server design, focusing on tool list composition, response payload optimization, and effective error handling.