The Amazon Web Services (AWS) Open Data Sponsorship Program has hosted the Common Crawl open repository of web data at no cost since January 2012. It has become one of the most important sources of training data for the large language models (LLMs) reshaping industries. This is the story of a 14-year collaboration between a small nonprofit and AWS, and how open data infrastructure became the foundation of the AI era.