The data black hole at the center of AI
Dwarkesh Patel argues that AI progress is hitting a "data black hole" where the most valuable training data (from expert domains, rare events, and physical interactions) is exponentially harder to obtain, potentially limiting the continued scaling of AI models even as compute becomes cheaper.
Background
- The "data black hole" is the idea that state-of-the-art AI models (like GPT-4) are already trained on almost all publicly available high-quality text on the internet. There is very little fresh, useful data left to improve them.
- Dwarkesh Patel is a tech analyst and writer who interviews AI researchers and publishes a popular newsletter on AI progress and its bottlenecks.
- This is the central supply-side dilemma for scaling AI: companies can throw more compute at models, but if they've already absorbed the entire useful web, performance gains from training on larger datasets will hit a wall.
- The piece connects to ongoing debates about whether "scaling laws" (bigger models + more data = better performance) are breaking down, and whether synthetic data or new data sources can fill the gap before companies hit this wall.