Data lineage for machine learning traces how information flows through ML pipelines from raw inputs through feature engineering and model training. The article argues that most ML production failures stem from upstream data quality issues rather than model degradation, and that fragmented lineage across multiple tools prevents teams from detecting these problems. DataHub’s approach provides full-stack metadata integration with column-level precision to enable root cause analysis, detect target leakage, and ensure reproducibility across the ML ecosystem.