Robots.txt for Beginners: What Every WordPress Site Owner Should Know
There’s a file on your WordPress site, probably under ten lines long, that can erase your entire website from Google with a single misplaced character. Most site owners never open it until something breaks. This is the guide to opening it before that happens.
What Is Robots.txt, Really?
Every respectable search engine crawler — Googlebot, Bingbot, and the newer AI crawlers like Google-Extended — knocks on the same door before it starts reading your site: yourdomain.com/robots.txt.
That file is a plain-text note that says which parts of your site crawlers may visit. A typical one looks like this:
User-agent: * Disallow: /wp-admin/
That’s the whole idea. It’s a request, not a lock. Well-behaved crawlers read it and comply. Scrapers, spam bots, and bad actors ignore it completely — so robots.txt is useless as a security tool, and you should never treat it as one.
One more thing people get wrong from the start: the filename has to be exactly robots.txt, lowercase, sitting at the root of your domain. A file at yourdomain.com/blog/robots.txt is invisible to crawlers. They only ever check the root.
What Robots.txt Cannot Do (This Is the Part That Matters Most)
Here’s the misunderstanding that causes real damage: robots.txt does not keep pages out of Google.
Google says this explicitly in its own robots.txt documentation — the file “is not a mechanism for keeping a web page out of Google.” All a Disallow rule does is ask crawlers not to visit a URL. If Google discovers that URL some other way (say, through a link from another site), it can still put it in the index. You’ll then see it in Search Console marked “Indexed, though blocked by robots.txt” — a ghost listing with no description, which is worse than useless.
So if your actual goal is keeping a page out of search results, you need one of these instead:
- A noindex directive (a meta tag or HTTP header), which tells search engines “don’t put this in your index.” This is the correct tool for thin pages, thank-you pages, internal search results, and staging content.
- Password protection, for anything that genuinely must stay private.
Robots.txt controls crawling. Noindex controls indexing. They are different jobs, and mixing them up is how sites end up half-invisible in search results for months. If you’re still getting your head around how search engines find, crawl, and rank pages in the first place, start with our overview of what digital marketing and SEO actually cover.
How to Read a Robots.txt File
You only need to understand four directives to read any robots.txt file ever written:
- User-agent: which crawler the rules apply to.
*means every crawler. A specific name likeGooglebottargets just that one. - Disallow: paths crawlers should not visit.
/wp-admin/blocks everything under that folder. - Allow: an exception to a Disallow. Used sparingly, and WordPress relies on one (more on that below).
- Sitemap: where your XML sitemap lives, so crawlers can find it without guessing.
- Lines starting with # are comments — ignored by crawlers, useful for humans.
(That’s the everyday vocabulary. Google’s full specification goes deeper, but you won’t need it for anything in this guide.)
Two combinations worth memorizing. This allows everything:
User-agent: * Disallow:
(A blank Disallow disallows nothing.) And this blocks everything:
User-agent: * Disallow: /
That second one is the nuclear option. If you ever see it on a live site, fix it immediately — it’s telling every search engine on the internet to stay out of your entire website.
A Safe Starter File for WordPress
For the vast majority of WordPress sites, this is all you need:
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://yoursite.com/sitemap_index.xml
Line by line:
Disallow: /wp-admin/keeps crawlers out of your login and dashboard area. There’s nothing there worth indexing, and it contains sensitive paths.Allow: /wp-admin/admin-ajax.phpis the exception. That file handles background requests on your site’s front end (forms, filters, live search), so blocking it can break how your pages behave for visitors and for crawlers alike.- The
Sitemap:line points crawlers at your XML sitemap. If you use Rank Math or Yoast, your sitemap lives at/sitemap_index.xml— swap in your real domain.
Honest disclosure: this is essentially what WordPress already generates for you. Out of the box, WordPress serves a “virtual” robots.txt (no physical file needed) containing the Disallow: /wp-admin/ and the admin-ajax exception. If you run Rank Math, it manages the virtual file for you and adds the sitemap line automatically. So if you’ve never touched robots.txt, you’re probably already fine — the mistakes below are what you need to watch for.
How to Edit Robots.txt in WordPress
Method 1: With Rank Math (recommended for this site)
Rank Math publishes its own guide to editing robots.txt, and the process is short:
- In your WordPress dashboard, go to Rank Math SEO → General Settings → Edit robots.txt.
- If you don’t see the “Edit robots.txt” option, Rank Math is running in Easy Mode. Switch it to Advanced Mode first (Rank Math SEO → Dashboard), then look again.
- Edit the rules and click Save Changes.
- Open
yourdomain.com/robots.txtin a private browser window and confirm your changes are live.
One gotcha: if there’s a physical robots.txt file sitting in your site’s root folder (usually public_html), it overrides everything Rank Math does — the editor may even appear greyed out. In that case, delete or edit the physical file directly through your hosting File Manager or FTP. Never run two editors at once; that’s how stale rules survive for years.
Method 2: Manually
- Open a plain text editor (Notepad, not Word).
- Write your rules and save the file as
robots.txt. - Upload it to your site’s root directory via FTP or your hosting control panel’s File Manager.
- Visit
yourdomain.com/robots.txtto confirm it’s serving.
5 Robots.txt Mistakes That Actually Hurt WordPress Sites
1. Blocking your CSS and JavaScript.
The most common bad advice on the internet is to add Disallow: /wp-includes/ or Disallow: /wp-content/. Don’t. Google needs your CSS and JS files to render pages properly — they live in exactly those folders. Block them and Google judges an unstyled, broken-looking version of your page. You’ll spot this in Search Console’s URL Inspection tool: page resources “could not be loaded” and the rendered screenshot looks nothing like your site.
2. Disallowing categories, tags, or author pages to “save crawl budget.”
Blocking archive pages in robots.txt doesn’t remove them from the index (see the “Indexed, though blocked” trap above) — and worse, it stops Google from reading any noindex tag you put on them. If you don’t want archives indexed, use noindex, follow instead. Blocking the crawl and the indexing signal fight each other; let the noindex do its job.
3. Leaving “Discourage search engines” switched on.
WordPress has a checkbox under Settings → Reading → Search engine visibility labeled “Discourage search engines from indexing this site.” It’s meant for staging sites. The trap: people build on a staging copy, migrate it live, and forget the box is ticked. Since WordPress 5.3, this setting adds a noindex tag across your site — your pages simply won’t appear in Google, and nothing in robots.txt will explain why. After every migration, check that box first.
4. Editing in your SEO plugin while a physical file exists.
As mentioned above, a physical robots.txt in your root folder wins over anything Rank Math or Yoast generates. You can edit the plugin screen all day and nothing will change on the live file. If your edits aren’t showing up at yourdomain.com/robots.txt, check for the physical file.
5. Using robots.txt to hide a page.
It bears repeating because it’s the single most expensive mistake: disallowing a URL does not deindex it. A disallowed page with inbound links can sit in Google’s index indefinitely as a title-only ghost result. If a page shouldn’t be in search results, noindex it or password-protect it — don’t just block the crawl.
How to Test Your Robots.txt
Don’t trust your edits — verify them:
- Look at the live file. Open
yourdomain.com/robots.txtin a private window. Confirm it’s the version you think it is. - Check Google’s view. In Search Console, go to Settings → robots.txt. This report shows the exact file Google fetched, the HTTP status, when it last fetched it, and any parsing errors. (It’s available for domain-level properties and root URL-prefix properties.)
- Test a specific URL. Paste any page URL into Search Console’s URL Inspection tool. If robots.txt is blocking it, the coverage section says so plainly: “Blocked by robots.txt.” Hit “Test Live URL” to check against your current file rather than Google’s cached copy.
Frequently Asked Questions
Is robots.txt still relevant?
Yes. Every major crawler still fetches it first, and it’s now the file AI crawlers check too — Google’s Google-Extended crawler, which governs AI training use of your content, is controlled through robots.txt rules like any other user-agent.
How can I create a robots.txt file?
Two ways: let your SEO plugin generate it (Rank Math → General Settings → Edit robots.txt, or Yoast → SEO → Tools → File editor), or create a plain-text file named robots.txt and upload it to your site’s root folder via FTP or File Manager. The manual file overrides the plugin’s version.
Is robots.txt enforceable?
No. It’s a voluntary convention — the polite crawlers obey it, the rude ones don’t. That’s also why you should never list genuinely sensitive paths in it: the file is public, so a Disallow: /secret-admin/ line is just advertising the path.
What should my robots.txt look like?
For most small WordPress sites, the safe starter file in this guide is the complete answer: block /wp-admin/, allow admin-ajax.php, and add your sitemap line. If you haven’t customized anything, WordPress and Rank Math already serve something very close to this — the best robots.txt is usually the one you barely need to think about.
Robots.txt is a small file with outsized consequences. Check yours after every migration, every staging push, and every time an SEO plugin changes hands on your site — it takes thirty seconds, and it’s the cheapest insurance in technical SEO.
Next in this series: what a sitemap actually is, and how to submit yours to Google Search Console the right way.