How Theta EdgeCloud is splitting AI workloads across seperate GPUs to keep large models running faster and more reliably. The Theta EdgeCloud team has completed a benchmark testing a more efficient way to serve large language models in production. LLM inference involves two phases of work, namely prefill and decode, that sit awkwardly together on the same hardware.