This post argues that agent memory requires more than larger context windows, as production agents must remember facts across hours, days, and sessions, update them when reality changes, and retrieve them reliably. Traditional benchmarks fall short because they fail to test whether an agent can maintain consistent beliefs about a user, task, or world state across time. BEAM is examined as a benchmark designed to address these gaps and better predict real-world memory system performance.