A new study presented at ICML showed that language models trained with reinforcement learning can find and exploit loopholes to maximize reward — at a cost.
How training environments can teach AI models to misbehave
calendar_today
July 9, 2026
domain
ibm