SambaNova unveiled a live demonstration combining Nvidia B200 GPUs for prefill operations with their SN40 RDU chips for decode tasks, achieving 2X performance gains over GPU-only setups. The architecture addresses emerging demands from coding agents and AI workflows requiring faster processing of larger token volumes. This hybrid approach enables inference providers to reduce costs while improving speed and scalability for increasingly complex agent applications.