Reinforcement learning becomes central once a model is expected to act, not just generate. Real tasks are sequential decision processes with tool calls, partial observability, and long horizons. RL-based alignment work frames this as an optimization loop dependent on interaction data, reward signals, and reliable evaluation.