The article uses the metaphor of a "data black hole" to describe how massive, invisible datasets underpin the visible capabilities of modern AI systems, pulling in vast amounts of information to train models whose inner workings remain opaque even as their outputs appear impressive.
Background
This post uses the "data black hole" metaphor to explain why AI scaling is hitting a wall. The core idea: the most capable large language models (LLMs) like GPT-4 and Claude have already consumed nearly all accessible, high-quality human-written text on the internet. This "training data" is the gravitational mass that gives AIs their capabilities. Frontier AI labs (OpenAI, Google DeepMind, Anthropic, Meta) spent 2022-2024 scaling up compute and data, but the supply of novel, clean text is finite. The author argues that future progress will increasingly depend on synthetic data (AI-generated text used to train other AIs) or reinforcement learning from human feedback (RLHF), neither of which is a perfect substitute for original human data. This context matters because it challenges the prevailing narrative that AI capabilities will continue improving at the same exponential rate — if the "fuel" of new data runs out, scaling could decelerate sharply. Dwarkesh Patel runs the Dwarkesh Podcast, which focuses on long-form interviews with AI researchers, and the post reflects debates inside the AI safety and scaling communities.
Dwarkesh Patel argues that AI progress is hitting a "data black hole" where the most valuable training data (from expert domains, rare events, and physical interactions) is exponentially harder to obtain, potentially limiting the continued scaling of AI models even as compute becomes cheaper.
This video explores the issue of data transparency in AI, highlighting how major AI companies often withhold details about the datasets used to train their models. It examines the consequences of this lack of transparency, including potential biases, legal challenges, and the difficulty of auditing AI systems for accountability.