Common Crawl publishes petabytes of web crawl data on S3. With DuckDB and MotherDuck you can query the Common Crawl dataset directly, no cluster and no download, and measure how fast the vibe-coded web is growing. |…
1 comment
dmkii3 days ago
You can basically SELECT * the internet by querying the public Common Crawl dataset on S3 with DuckDB. It's a great way to explore how the AI-pilled web (Vercel, Lovable, Replit, Bolt, Cloudflare, GitHub) compare to the traditional Wordpress-Blogspot hegemony.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 5 days ago
- Hacker News · 1 points · 5 days ago
- Show HN: Last Internet Connectiongithub.comHacker News · 1 points · 8 days ago
- Hacker News · 2 points · 7 days ago
- Be Careful with Your Select * Queriesnotesonsystems.comHacker News · 1 points · 11 days ago
- Hacker News · 60 points · 13 days ago