Advanced fusion kernels improve training throughput for mixture-of-experts (MoE) models, which enable larger model capacity while activating only a subset of parameters for each token for improved performance scaling.
Need help?
Contact usAdvanced fusion kernels improve training throughput for mixture-of-experts (MoE) models, which enable larger model capacity while activating only a subset of parameters for each token for improved performance scaling.