The email arrives on a Thursday night. Subject line in bold, tone set to full siren. “New reason preventing your pages from being indexed.” A client reads it on their phone, heart rate climbing, and forwards it to us with three words. “Is this bad?”
It isn’t bad. It isn’t anything. The site is healthy, the file in question is doing its job with surgical precision, and Google is reporting the obvious back to a person who has no way of knowing it’s obvious. That gap, between what the machine says and what a human hears, is the whole story. So let’s tell it properly.
Google Search Console · Automated Notification
New reason preventing your pages from being indexed
Search Console has identified that some pages on your site are not being indexed due to the following new reason.
If this reason is not intentional, we recommend that you fix it in order to get affected pages indexed and appearing on Google.
Blocked by robots.txtRead that last line again. “If this reason is not intentional.” Google does not know whether it was intentional. It cannot tell the difference between a professional who blocked a staging folder on purpose and a panicked amateur who just deleted their own traffic. So it defaults to alarm and hands the interpretation to the one person least equipped to do it. The site owner.
The problem that is not a problem
“Blocked by robots.txt” is not an error. It is Google confirming that a rule you wrote is working exactly as written. A crawler asked to visit a URL, your file told it not to, and the crawler obeyed. That is the entire event. It is the digital equivalent of your alarm company calling to report that your locked door is, in fact, locked.
A locked door is not a break-in. It is the lock doing its job.
When that notice fires on a WEBPRO-built site, the flagged URLs are precisely the ones we intended to keep out of the index. Admin folders. Staging directories. Backup archives. Raw database files. The plumbing of the site that has no business showing up in a search result. Google found a link to one of them, respected the instruction not to crawl it, and then wrote us a letter about it in a font that says emergency.
What a masterfully built file actually looks like
Most agencies treat robots.txt as an afterthought. A three-line file copied from a forum in 2014, never touched again. Here is a portion of what we build instead, and every line is a deliberate decision.
robots.txt - Smile Worthy Orthodontics - built by WEBPRO
Look at what is happening here, because almost no one builds to this standard for what people dismiss as a simple text file.
- 01
Content signals with a legal spine
The Content-Signal directives are backed by an express reservation of rights under Article 4 of EU Directive 2019/790. This is not a suggestion to a scraper. It is a documented legal position on how the content may and may not be used.
- 02
AI scrapers excluded by name
GPTBot, Google-Extended, CCBot, Bytespider, and more, each shut out individually. Search crawlers stay welcome so the site ranks. Training crawlers get nothing. That is a choice about who profits from the client’s words.
- 03
WordPress hardened, not exposed
Admin, staging, backups, and raw database and archive files are held out of the index. This is the exact rule set that triggered the scary email, and it is exactly what a competent build is supposed to do.
- 04
Sitemap declared, search invited in
The front door is wide open. The sitemap is announced, every public page is crawlable, and the result speaks for itself. This site ranks and performs. Nothing is broken. Nothing is even close to broken.
So we have a file that is more thorough, more current, and more legally aware than what 95% of the industry ships. And the reward for that rigor is an automated email telling the client something might be wrong. Which brings us to the question nobody at Google wants asked out loud.
Why the alarm is built to sound like an alarm
Nobody at Google sat down and wrote a scare email about a 31-year veteran who was solicited by Google, who knows their rules cold, and who built the file to a higher standard than their own documentation demands. That is not the mechanism. The mechanism is quieter and more honest than a conspiracy, and far more effective.
Google’s revenue does not come from organic search. It comes from paid search. Every dollar of it. So the entire system is tuned, not by malice but by incentive, to make organic feel uncertain, fragile, and frightening. A confident site owner is a problem. A nervous one reaches for the ad account.
First they scare you about organic. Then they sell you the way out. It is called PPC, and the door only swings one direction.
Read the email again with that lens. It does not say “we noticed an intentional block, all good.” It says “preventing your pages,” “not being indexed,” “we recommend you fix it.” Every word is chosen to raise the pulse of someone who does not know better. And a frightened owner who does not have a WEBPRO in their corner does the predictable thing. They panic, they second-guess a perfectly good site, and eventually they buy their way back to a visibility they never actually lost. Straight into paid purgatory, confidence handed over at the door.
That is the real message hiding inside a routine notification. Manufacture just enough doubt about the free channel to make the paid channel look like safety.