Qwen3.8-2.4T-A95B, Alibaba’s 2.4T-parameter open-weights MoE model, is now on DigitalOcean Serverless Inference at $2/$6 per 1M tokens, served with NVFP4-quantized weights on NVIDIA HGX B300. This post covers measured TTFT and throughput, quickstart code for tool calling and structured outputs, honest benchmark tradeoffs, and when to route work to it.