TL;DR I fine-tuned a small multimodal model for Mux-specific video intelligence workflows, like transcript-based summaries and chapter generation. I then integrated that model into the open-source @mux/ai SDK . Along the way, I added Baseten as a provider, generated 10,000 synthetic JSONL training examples, and used LoRA to fine-tune Mistral Small 3.1.
Fine-tuning a multimodal model for video intelligence
calendar_today
May 28, 2026
domain
mux-com