Self-supervised models increasingly convert medical images into quantitative phenotypes for biological discovery, but statistical reproducibility does not establish that a learned phenotype represents the intended anatomy. We trained a video masked-autoencoder on 69,932 UK Biobank cardiac cine-MRI studies and performed genome-wide association analysis of its latent representation. Although 18 of 20 leading axes were heritable with well-calibrated statistics, the representation encoded substantial field-of-view information: body size, stature and imaging centre (linear-probe R^2=0.55 for site);