Prime Intellect released version 0.6.0 of prime-rl, enabling training of trillion-parameter scale models on heavy agentic workloads at high efficiency. The system trains GLM-5 on software engineering tasks at 131k sequence length with sub-5-minute step times using only 28 H200 nodes, through optimizations spanning low-precision inference, prefill/decode separation, and disaggregated trainer-inference architectures.