Troubleshooting a distributed application means correlating signals across alarms, traces, logs, and deployments, usually across several consoles while the clock is running. Imagine your on-call engineer gets paged at 2 AM. P99 latency on the checkout API has spiked past the SLO threshold.