Mengqi Yuan (XLANG Lab, University of Hong Kong) presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows where even frontier AI agents complete only 20.6% of tasks outright after 300+ steps each. The post OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks appeared first on Snorkel AI .