Kakao describes the engineering challenges of serving Kanana-O, a multimodal model that understands text, images, and audio and responds with natural text and speech, in a real-time voice conversation service, and the optimizations behind the Kanana-Omni serving server.