We are releasing LFM2-8B-A1B, our first on-device Mixture-of-Experts (MoE) with 8.3B total parameters and 1.5B active parameters per token. By activating only a sparse subset of parameters during inference, LFM2-8B-A1B delivers larger model quality with the compute of a 1.5B-class model. It trades a modest increase in memory footprint for higher quality and speed compared to dense models. This enables fast, private, low-latency applications on modern phones, tablets, and laptops.