Research · scanned September 25, 2026

1 in 12 YC startups show ChatGPT's crawler an empty homepage

We ran our free crawler check on the homepage of every active Y Combinator company, 4,333 sites, to see what the crawlers behind ChatGPT search, Claude and Perplexity actually receive. For 8.4% of the homepages we could check, the answer was a page with almost no words on it.

Check your own site

The same check we ran on every homepage below. Free, no signup, takes a few seconds.

The numbers

8.4%

serve AI crawlers an empty page

318 of 3,795 homepages

1.3%

block an AI search crawler

49 of 3,824 homepages

9.5%

have one problem or the other

362 of 3,824 homepages

Blocking gets the attention, but it's rare: most YC homepages let every AI search crawler in. The common failure is quieter. The crawler arrives, the server answers, and the page it gets back is an empty shell that a browser fills in with JavaScript. The crawlers behind AI search don't run JavaScript, so that shell is all they have to go on.

Newer companies are more likely to ship one

Homepages that are empty to AI crawlers, by YC batch

Share of checked homepages in each group.

2025–26 batches10.2% (122 of 1,193)
2022–24 batches9.2% (113 of 1,233)
Before 20226.1% (83 of 1,368)

The newest batches have the highest rate, and they're the companies with the most to gain from turning up when someone asks an AI for a tool like theirs.

What an empty homepage looks like to a crawler

A typical one, trimmed. This is the whole page as a crawler receives it: a title, a script tag, and an empty element for JavaScript to fill.

<!doctype html>
<html>
  <head>
    <title>Acme – The platform for …</title>
    <script type="module" src="/assets/index-3f9a1c.js"></script>
  </head>
  <body>
    <div id="root"></div>
  </body>
</html>

It isn't one framework's fault. Of the 318 empty homepages, 190 were built with React, 54 with Next.js (which renders on the server by default, so these had turned that off somewhere), 13 with Vue, and 55 with a setup we couldn't identify. Hosting doesn't explain it either: they were spread across Cloudflare, Vercel, CloudFront and others. It's a build setting, and it's easy to miss because the site looks fine in a browser.

Why Google doesn't catch it

Google renders JavaScript before indexing, so these sites can rank normally in Google and look fine in every SEO tool that checks Google. ChatGPT search, Claude and Perplexity read the HTML as it arrives. When that HTML is empty, there's nothing for them to quote, so they cite someone else.

How to fix it

Put the content in the first HTML response. Turn on server-side rendering or static prerendering for the pages you want cited. On React with Vite or Create React App, prerender the marketing pages at build time. On Next.js, render the page's content on the server rather than only in client components. Nothing needs rewriting. Then confirm it from a terminal:

curl -s -A "OAI-SearchBot" https://yoursite.com/ | grep -c "<h1"

0 means crawlers still get an empty page. 1 or more means your heading is in the HTML.

On blocking

49 homepages (1.3%) block at least one AI search crawler (OAI-SearchBot, Claude-SearchBot or PerplexityBot), in robots.txt or at the server. 14 refuse every AI crawler we sent. Separately, 6.8% block a training crawler (GPTBot, ClaudeBot or Google-Extended). That's a reasonable choice and it doesn't stop a site being cited, because the search crawlers are separate bots.

Method

  • Population: every company listed as active or public in the community-maintained YC directory export (yc-oss), one homepage per host: 4,333 sites, scanned on 2026-09-25.
  • Each homepage got the check on this page, unchanged: robots.txt evaluated for each crawler, one request per crawler user-agent, and one plain fetch analysed for content that only appears with JavaScript.
  • "Empty" means the HTML carries fewer than 150 words of text and is built to be filled in by JavaScript (an empty app element, a script-only body, a notice asking visitors to turn JavaScript on). A page that already serves more text than that isn't counted, whatever else it contains.
  • Every homepage flagged as empty was fetched again as OAI-SearchBot. 10 of them serve bots full HTML (dynamic rendering) and are not counted as empty.
  • 509 homepages couldn't be checked and are left out of every percentage: 331 timed out or gave no clear answer, 139 refused an ordinary browser request as well, 34 were missing, and 5 errored.
  • Server blocks are tested by sending each crawler's name from our own network. A server that admits the real crawler only from its vendor's addresses would look blocked to us, so the blocking figures are an upper bound.
  • Only homepages were checked. We don't name companies, and we aren't affiliated with Y Combinator.