At Malt, the Machine Learning Engineering (MLE) team is responsible for building models, making them available in production and monitoring them. With over 20 FastAPI HTTP endpoints in production (our Matching and Recommender systems) and 40+ daily Airflow jobs feeding our medallion data lake, “keeping the lights on” is a high-stakes task. For a long time, our bi-daily routine felt like an archaeological dig.