Hi Fernando — this is genuinely useful, thanks for going through the repo and reporting back in such detail. On your direct question : no, I haven’t gotten a16 through the attention block either. My own a16_w16 usage is limited to input_layer1 (the embedding input, far from attention — matches what you saw) and ew_add* layers on the encoder recipe.