With the advance of Large Language Models and Vision Models, many empirical results have proven that model performance is positively correlated with model capacity (number of parameters) and data quantity (number of training samples). Such correlations have been modeled by the so-called scaling law . But despite its success in vision models and large language models, scaling law’s generalizability to ML self-driving models has been unknown.