Machine learning (ML) models used for inference on Kubernetes are often several gigabytes in size. When these models are embedded in container images, images become oversized and pod scheduling slows. More critically, inference pods are inherently stateful.