Amazon SageMaker AI now emits over 100 detailed inference metrics for monitoring LLM endpoints, and the SageMaker Insights dashboard in CloudWatch provides observability through Performance, Capacity, and Reliability tabs tracking GPU health, token-level latency, KV cache pressure, and AZ traffic distribution. Features include per-accelerator utilization metrics, cold-start diagnostics, insufficient capacity error tracking, and PromQL integration for connecting external tools like Grafana.