Skip to content

Home / Notes / Technical SEO

Work notes ·

Robots.txt and noindex solve different problems

Why a crawl block can get in the way of removing a page from search, and how I check WordPress search visibility from the outside.

By Shane Rounce · Published · 2 min read

A crawl instruction and an indexing instruction sound similar enough to be confusing. They act at different points.

In September 2026 I tightened the way my SEO tooling respected WordPress’s native search-visibility setting. I also added checks for the pieces around it: the served robots file, page output, response headers and whether the site was actually responding properly.

The short version

robots.txt tells compliant crawlers which requests they may make. A noindex instruction tells a supporting search engine not to retain a page in its index once it has fetched and read that instruction.

If a crawl block prevents that fetch, the search engine cannot see the page’s new noindex. Google spells this out in its noindex guidance. That is why an already indexed test page needs a different removal plan from a new test site that has never been public.

Use the setting the site already has

WordPress already has a search-visibility control. I changed the plugin to respect it throughout its output rather than introducing another competing definition of whether the site was public.

For a deliberate de-indexing process, the tooling gained a separate mode that allowed the relevant crawl while serving the removal instruction. That needs a clear owner and follow-up checks. It is not a reason to leave a private development site accessible.

Neither robots rules nor noindex provide access control. If information must stay private, I put authentication or an appropriate access restriction in front of it.

Read the response, not just the checkbox

A physical robots file can be served before WordPress gets involved. A page cache can return HTML generated before the setting changed. The admin screen can therefore be correct while the public response is still wrong.

  • Fetch the actual robots file at the site root.
  • Inspect a public page’s response headers and HTML.
  • Check a second page type so a template-specific rule is not missed.
  • Confirm the response is a real working page, not a server error dressed up as a passed check.
  • Repeat after clearing the relevant cache and after deployment.

I keep the intended behaviour written down for each environment. A test site and a public site often need different outcomes, and a deployment should not silently swap one for the other.

The useful result is agreement between the setting, the served instructions and the purpose of the site. A green tick beside only one of those is not enough.

Notes from my own work, with client details left out. Published as a retrospective on 6 October 2026.

Got something similar to untangle?

I help with SEO, from finding the problem to making the change. Tell me what you are trying to improve.

Email Shane