How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Larger RL batches only help when throughput beats the sample penalty

calendar_today September 23, 2026 person ZiniuLi domain aligned-news

Ziniu Li highlights a study finding that learning-rate retuning is the difference between a 29% GRPO speedup and a 42% slowdown. The underlying preprint was first submitted in August.

open_in_new Read original post