Free SEO tool

Googlebot 2MB Checker

Googlebot only reads the first 2MB of your HTML. Paste a URL to see exactly what Google gets: the raw HTML up to the cutoff, and a rendered screenshot of the page as it looks at that point.

Crawl as

Why Google's 2MB limit matters

Google's documentation for Googlebot says it plainly:

"Googlebot crawls the first 2MB of a supported file type, and the first 64MB of a PDF file."
(Google Search Central: Googlebot)

Once Googlebot hits the cutoff, it stops downloading and sends only what it already has on for indexing. Content, internal links, structured data or a canonical tag that sit past the 2MB mark might as well not exist.

Most pages are far below 2MB. The ones that go over usually have huge inline JavaScript bundles (framework hydration state such as __NEXT_DATA__), base64-encoded images or fonts, big inline SVG sprites, or very long product listings rendered on a single page. This checker shows how big your HTML is, where the cutoff falls, which important tags are past it, and what's using up the space.

How this checker works

  1. We fetch your URL as Googlebot (or as a regular Chrome user, if you pick that), following redirects, and measure the uncompressed HTML, which is what Google measures.
  2. We cut the HTML at 2MB and show you that raw HTML exactly as Google received it.
  3. We load that cut-off HTML in a headless Chrome browser, applying the same 2MB cap to every CSS and JavaScript file, and take a screenshot.

We don't store the pages, the HTML, or the screenshots.

FAQ

Does Google really only read the first 2MB of a page?

Yes. Google's Googlebot documentation says Googlebot crawls the first 2MB of a supported file type (and the first 64MB of a PDF). Anything after the cutoff isn't used for indexing.

Does gzip or Brotli compression help?

No. The limit applies to the uncompressed data. A page that transfers as 400KB gzipped but expands to 3MB of HTML is still over the limit.

Do my CSS and JavaScript files count toward the 2MB?

External files don't. Each resource is fetched separately and has its own 2MB limit. Inline <script> and <style> blocks, inline SVG and base64 data URIs are part of the HTML, so they do count.

My page is over 2MB. What should I fix first?

Make sure your <head> tags (title, meta description, canonical, hreflang, robots) and main content come early in the HTML. Then shrink the biggest items in the "What's using your 2MB" breakdown: move inline scripts and styles into external files, replace base64 images with normal image URLs, and paginate or lazy-load very long lists.

Why does my page show a "blocked" warning?

Some sites use bot protection (Cloudflare, Akamai, DataDome and others) that challenges automated requests from cloud servers. Many of these block any request that claims to be Googlebot but doesn't come from Google's own IP addresses, which is how they catch fake Googlebots. The real Googlebot is verified by its IP address, so it usually gets your actual page. Try crawling as "User Smartphone" or "User Desktop"; if the site still blocks us, its protection is based on our server's IP address.

Is the screenshot exactly what Google renders?

It's a close approximation. We use a recent headless Chrome with Googlebot's user agent and viewport, but Google's Web Rendering Service has its own caching, timeouts and resource rules. For an official view of a single URL, use the URL Inspection tool in Google Search Console.

Want SEO that drives revenue, not just traffic?

Revenomics runs Full-Funnel SEO™ for brands that care about conversions.

Talk to Revenomics