This article examines how infrastructure choices significantly impact reinforcement learning training for coding agents. The research demonstrates that execution environment latency - from sandbox provisioning to command execution - compounds dramatically at scale, with differences ranging from 36 to 449 aggregate worker-hours across five infrastructure archetypes. The piece argues that optimizing the execution substrate is now critical for scaling coding-agent RL effectively.