JetBrains Research developed Step Rejection Fine-Tuning to improve LLM agent training by learning from failed attempts. The technique masks harmful steps while preserving valuable correct behaviors from incomplete trajectories.
Need help?
Contact usJetBrains Research developed Step Rejection Fine-Tuning to improve LLM agent training by learning from failed attempts. The technique masks harmful steps while preserving valuable correct behaviors from incomplete trajectories.