Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

What GPTBot sees before your React app hydrates

When GPTBot crawls a React site before hydration, it only sees the initial static HTML—empty containers, placeholder text, and skeleton loaders. This means the bot may miss dynamic content like user-specific data or client-rendered elements unless you implement SSR, pre-rendering, or structured data to make the page crawlable.

Background

- GPTBot is OpenAI's web crawler that scrapes public websites to train AI models like GPT-4 and ChatGPT. - "Hydration" in React refers to the process where a static server-rendered HTML page becomes interactive JavaScript client-side. Before hydration, the page is just raw HTML. - This article examines what GPTBot "sees" if it crawls a React site before hydration occurs — meaning it might encounter placeholder content (e.g., loading spinners, empty divs, "JavaScript required" messages) instead of the actual content that only appears after the app loads in a browser. - This matters because if GPTBot indexes incomplete or placeholder content, that can affect AI training data quality (bad signals for the model) and SEO (search engines might not see the real content either). - The piece is relevant to anyone building React-heavy web apps, especially given growing concern over how AI crawlers interact with modern JavaScript frameworks.

Related stories