GSC robots.txt Tester or Page Indexing: Which Check?

Robots.txt testing shows crawl permission on paths; Page indexing shows why URLs are in or out of search. Route to the right GSC check—SEO topic, not canonical troubleshooting.

GSC robots.txt Tester or Page Indexing: Which Check?

Search Console bundles crawl and indexing tools, so a worried blogger often opens whichever name sounds scariest. Robots.txt testing (and related robots reports) answers: under today’s robots.txt file, would Googlebot be allowed to request this path? Page indexing answers: which URLs are indexed or not indexed, and which reason bucket explains exclusions?

Those checks overlap in one narrow case—blocked by robots.txt in Page indexing—but they are not the same job. Passing a robots test does not prove posts rank. A not indexed URL does not always mean your Disallow lines are wrong.

This page is SEO routing for small blogs. It is not a canonical-tag troubleshooting guide. It is not Page indexing or Sitemaps: which report—that neighbor splits discovery file versus indexing ledger; this one splits crawl permission testing versus indexing outcomes.

Early owners:

Edited robots.txt or migration crawl fear → robots tester owner. URLs stuck not indexed with labels → Page indexing owner. One URL after a fix → URL Inspection. Wide multi-URL recovery → indexing pillar.

Official Help: Test robots.txt markup · How Google interprets robots.txt · Page indexing report · URL Inspection.

Disclosure: Educational GSC routing. No ranking guarantees or invented crawl timelines.

Table of contents

  1. Crawl permission versus indexing ledger
  2. Start with robots.txt testing when
  3. Start with Page indexing when
  4. When both reports point at the same fix
  5. Traffic drops that need a third screen
  6. Scenario: plugin Disallow on /blog/
  7. Mistakes this router targets
  8. FAQ
  9. Run one check, then open depth

Crawl permission versus indexing ledger

Robots.txt is a root-level file that tells compliant crawlers which paths they should not request. Testing in Search Console (or reading robots state in URL Inspection) shows rule matching for a sample path and user-agent—typically Googlebot—not whether the URL earns clicks.

Page indexing is a property-wide view of indexing states: indexed counts, not indexed counts, and sample URLs grouped by reasons such as Blocked by robots.txt, Crawled – currently not indexed, Duplicate, Google chose different canonical than user, or Not found (404).

Mechanism for confusion: both touch “Google cannot see my site.” Robots blocks fetch attempts for disallowed paths in the common case. Indexing reasons include robots plus noindex tags, redirect chains, soft 404s, and quality choices that never mention robots.txt.

If I were helping a friend after they installed a caching plugin, I would load https://example.com/robots.txt in a browser, compare to the robots.txt owner checklist, run one representative /blog/ path through robots testing, then open Page indexing only if URLs still sit in not indexed buckets after deploy.

Start with robots.txt testing when

Open robots.txt tester basics when:

  • You edited Disallow or Allow lines on the live file
  • You migrated hosts, staging rules leaked, or a plugin appended Disallow: /
  • You need to confirm sitemap URLs are not accidentally disallowed (Sitemaps owner pairs here)
  • Page indexing shows Blocked by robots.txt but you have not verified the live file yet

Why this check comes first in those cases: if Googlebot cannot fetch content because of a broad Disallow, fixing indexing requests is premature. Change production robots.txt on the server—Console tools interpret; they do not replace the file.

Test representative paths after changes: homepage, /blog/, one post, sitemap location. You rarely need four hundred post slugs in the tester; fix patterns, not individual permalinks.

Start with Page indexing when

Open Page indexing report basics when:

  • Important URLs remain not indexed with a labeled reason after robots looks clean
  • New posts stall in Discovered – currently not indexed or similar states
  • You need to prioritize which exclusion reason to fix across many URLs
  • You fixed noindex, canonical, or 404 issues and want to see whether bucket counts move

Page indexing teaches reason vocabulary—the ledger Google shows at property scale. It is the right first open when the panic is “my posts are not in Google,” not “I just pasted Disallow: /wp-admin/.”

For discovery-file confusion, add Page indexing versus Sitemaps after you pick indexing versus robots jobs.

When both reports point at the same fix

When Page indexing lists Blocked by robots.txt on sample URLs, robots testing should reproduce disallow on those paths unless cache or host mismatch interferes. The fix is still edit live robots.txt, deploy, verify in browser, retest sample paths, then watch Page indexing over days—not spam “Request indexing” on every post while /blog/ remains disallowed.

After the block clears, some URLs may move to Crawled – currently not indexed—a different bucket with different fixes documented on the Page indexing owner and why Google isn’t indexing pillar.

Traffic drops that need a third screen

Neither robots testing nor Page indexing alone explains every Performance dip. Impressions and clicks live under Search results → Performance (Coverage versus Performance when dashboards blur).

Manual spam penalties are yet another lane—Manual actions versus Page indexing when the worry is penalty language, not robots lines.

Canonical and duplicate URL work stays on dedicated canonical owners—this router is intentionally non-canonical per CashPilot cluster rules.

Scenario: plugin Disallow on /blog/

A security plugin adds Disallow: /blog/ to production robots.txt. Page indexing climbs for Blocked by robots.txt. The robots tester shows disallow on a sample post path. Fix: remove the bad line on the server, confirm live file, retest /blog/ sample, then monitor Page indexing—not rewrite meta descriptions on fifty posts.

If robots allows fetch but URLs stay not indexed, shift to Page indexing reason rows and URL Inspection on one representative URL.

Mistakes this router targets

  • Requesting indexing on hundreds of URLs while robots still disallows the directory
  • Treating a green robots test on / as proof every post is indexed
  • Using Page indexing alone to debug a fresh robots edit without reading the live file
  • Confusing this router with sitemap submission work or canonical tag audits

FAQ

Structured answers live in frontmatter faqs for schema; use them when you need quick disambiguation.

Run one check, then open depth

Crawl permission file → robots.txt tester basics plus robots.txt for small blogs. Indexing reasons → Page indexing report basics. Sitemap versus ledger → Page indexing or Sitemaps. Broad recovery → why Google isn’t indexing your pages. This URL picks which Console check matches your sentence so robots panic does not replace indexing literacy—or the reverse.

Keep learning

More guides in the same topic lane.

Make Money Online6 min read

Word to PDF or Compress PDF: Which Job?

Word to PDF exports DOCX to a fixed PDF; Compress PDF shrinks a PDF you already have. Match the job to Word source files versus overweight PDF attachments.

Make Money Online5 min read

Word OCR or PDF OCR: Which Job?

Word OCR recovers text from image-heavy DOCX files; PDF OCR adds search layers to scan PDFs. Choose by whether the client keeps Word or stays in PDF.

Make Money Online5 min read

PDF to PPTX or Protect PDF: Which Job?

PDF to PPTX rebuilds slides for editing in PowerPoint; Protect PDF adds a password to a finished PDF. Pick by whether they need slides or a locked PDF packet.