Common Crawl publishes petabytes of web crawl data on S3. With DuckDB and MotherDuck you can query the Common Crawl dataset directly, no cluster and no download, and measure how fast the vibe-coded web is growing.
Querying the entire internet (100 billion rows!) with MotherDuck
calendar_today
September 22, 2026
domain
motherduck