lintpageincident.log
checksenvironmentstoolsblogfaq
sign inget started
~/blog/robots-txt-testing-guide
SEORobots.txtTesting

robots.txt Testing Guide: How to Test Before You Deploy

Marius Orzaru·March 27, 2026·9 min read·updated August 29, 2026

Deploy first, test later is not a strategy

Your robots.txt is two lines of text that can make your entire site invisible to Google. Yet most teams treat it like a config file that doesn't need testing - write it once, deploy it, and forget it.

The problem is that robots.txt mistakes are silent. There's no build error, no console warning, no failing test. You only find out something's wrong when your traffic drops weeks later. This guide covers the whole loop: what the file actually controls, the syntax that trips people up, how to test it locally and in CI, the four ways it breaks in production, and how to verify the live version in seconds with a robots.txt validator.

What robots.txt does (and doesn't do)

A common misconception: robots.txt controls indexing. It doesn't. It controls crawling - whether a bot requests the URL at all. A page blocked by robots.txt can still appear in Google's index if other pages link to it. It just won't have any content in the snippet, because Google was never allowed to fetch it.

This distinction matters when you're trying to remove something from search results:

  • To stop a page being crawled, use robots.txt. Good for crawl budget, API routes, and admin areas.
  • To stop a page being indexed, use a noindex meta tag or X-Robots-Tag header. The page must remain crawlable for Google to see the directive.

Blocking a URL in robots.txt and adding noindex is a classic own-goal: Google can't crawl the page, so it never sees the noindex, so the URL can linger in the index indefinitely. Pick one.

The syntax you need to get right

A robots.txt file lives at the root of your domain (https://example.com/robots.txt) and follows a simple format:

User-agent: *
Allow: /
Disallow: /api/
Disallow: /admin/
Disallow: /auth/

Sitemap: https://example.com/sitemap.xml
  • User-agent specifies which crawler the rules apply to. * means all crawlers.
  • Allow explicitly permits crawling of a path.
  • Disallow tells crawlers not to request URLs matching the path.
  • Sitemap points crawlers to your XML sitemap. It is technically optional, but omitting it is a free miss - include it, and make sure the sitemap it points at actually validates.

Rules are matched by prefix, not by path segment. Disallow: /api/ blocks /api/scan, /api/users, and everything beneath it. That prefix behaviour is the source of most of the mistakes below.

Test locally before deploying

Serve your robots.txt locally

If you're using a static robots.txt in your public/ directory, you can inspect it directly. But if you're generating it dynamically (like with Next.js robots.ts), you need to actually serve it:

# Start your dev server
pnpm dev

# Fetch the generated robots.txt
curl http://localhost:3000/robots.txt

Compare the output against what you expect. The most critical thing to verify: your production robots.txt does NOT contain Disallow: /.

Check environment-specific logic

Many frameworks generate different robots.txt files for staging and production. This is the number one source of robots.txt disasters - the staging config leaks into production. If your robots.txt is dynamic, test both environments:

// Common pattern in Next.js robots.ts
import type { MetadataRoute } from 'next';

export default function robots(): MetadataRoute.Robots {
  const isProduction = process.env.NODE_ENV === 'production';

  return {
    rules: [
      {
        userAgent: '*',
        allow: '/',
        // This is where mistakes happen
        disallow: isProduction ? ['/api/', '/auth/'] : ['/'],
      },
    ],
    sitemap: `${process.env.NEXT_PUBLIC_SITE_URL}/sitemap.xml`,
  };
}

Test with NODE_ENV=production to verify the production output is correct.

Common syntax pitfalls

Robots.txt syntax is deceptively simple, but small mistakes have big consequences.

Typos in directives

Disallow has one "s." Write Dissallow and the rule is silently ignored - crawlers treat unrecognised directives as comments. Same goes for User-Agent vs User-agent (case matters for some crawlers).

There is no error for this. The file parses, returns 200, and quietly does nothing.

Missing trailing slashes

Disallow: /api matches /api, /api/scan, and also /api-docs. If you only meant to block the API directory, use Disallow: /api/ with a trailing slash.

# Blocks /api AND /api-docs (probably not what you want)
Disallow: /api

# Blocks only /api/ and its children
Disallow: /api/

Wildcard gotchas

Googlebot supports * wildcards, but not all crawlers do. And wildcards can be broader than you expect:

# This blocks any URL containing "admin" anywhere
Disallow: /*admin*

# This is probably what you meant
Disallow: /admin/

Conflicting rules

When Allow and Disallow conflict, the more specific rule wins. But "more specific" means the longer path, which isn't always intuitive:

User-agent: *
Disallow: /docs/
Allow: /docs/public/

# /docs/public/guide.html -> ALLOWED (more specific rule wins)
# /docs/internal/spec.html -> BLOCKED

Blocking CSS and JavaScript

Some robots.txt files block /assets/, /static/, or /_next/. This prevents Googlebot from rendering your pages, which means it can't properly evaluate your content:

# Don't do this
User-agent: *
Disallow: /_next/
Disallow: /static/

Google needs your CSS and JavaScript to render pages like a browser. Blocking those resources can hurt rankings because Google can't see what the page actually looks like.

The four ways this breaks in production

Every scenario below passes local testing. They break at deploy time or later.

1. The staging-to-production leak

Your staging environment has a restrictive robots.txt to keep test content out of Google, which is correct for staging. The disaster happens when your pipeline copies the build output - including robots.txt - to production without environment-specific overrides. We have seen this with:

  • Docker builds that bake robots.txt into the image at build time, using the staging config
  • CI/CD pipelines that run a build with staging environment variables, then deploy that output to production
  • Monorepos where a shared public/robots.txt is used across environments without conditional logic
  • Platform migrations where the new host auto-generates a restrictive default that overrides your file

The fix isn't "remember to change it" - humans forget. The fix is the CI check further down this page.

2. The CDN cache that won't let go

You fix robots.txt, deploy, and verify it at the origin. But your CDN is still serving the old restrictive version. Google keeps seeing Disallow: / for hours, or days if the TTL is long.

# Correct at origin
curl -H "Cache-Control: no-cache" https://yourdomain.com/robots.txt

# But the edge may still be serving the old file - check the age
curl -sI https://yourdomain.com/robots.txt | grep -i "age\|cache-control"

After fixing a robots.txt problem, purge the CDN cache for that specific file. Vercel handles this on deploy but is worth verifying; Cloudflare and CloudFront need an explicit purge or invalidation for /robots.txt.

3. The framework upgrade that regenerates the file

You upgrade a framework or CMS and the new version writes a robots.txt with different defaults, overwriting your configuration. Common with Next.js robots.ts during major version upgrades, WordPress updates that regenerate virtual rules, and headless CMS builds that generate the file from a template.

If your robots.txt is generated dynamically, add a test that asserts the output. If it's static, make sure the build can't overwrite it.

4. The domain migration nobody told SEO about

Your team moves from www.example.com to example.com. DNS is updated, redirects are in place, the site works. But robots.txt on the new canonical domain either 404s or still carries the old restrictive rules.

Google treats robots.txt per domain. A correct file on www.example.com does nothing for example.com. After any domain change, verify the file on the new host.

Why standard monitoring misses all of this

Robots.txt failures are uniquely hard to detect:

  • No errors in your logs. The file returns 200. It's valid. It just says the wrong thing.
  • No impact on user experience. Every visitor browses normally. Only bots are affected.
  • Delayed symptoms. Google doesn't de-index instantly. Traffic declines over days, which makes it hard to correlate with a specific deploy.
  • No alerts from uptime monitoring. Your checks confirm the site is up, not that robots.txt says the right thing.

By the time someone notices the drop, investigates, finds the crawl block, fixes it, and waits for a re-crawl, weeks of organic traffic are gone. On a low-authority domain, recovery takes just as long again.

Test with Google Search Console

Search Console has a robots.txt report under Settings > robots.txt. It shows whether your file is accessible, any syntax warnings, and which version Google last fetched.

This is the authoritative test because it uses Google's own parser. If Search Console says a URL is blocked, that's what Googlebot will do.

The limitation: it only works for verified properties, and it only tests the live production file. It can't tell you whether the file you're about to deploy is safe.

Validate programmatically in CI

This is the check that actually prevents the staging leak. Run it against a preview deployment before promoting to production:

# Build and serve, then assert the output
ROBOTS=$(curl -s "$DEPLOY_URL/robots.txt")

if echo "$ROBOTS" | grep -qE "^Disallow: /$"; then
  echo "ERROR: robots.txt blocks all crawling"
  exit 1
fi

if ! echo "$ROBOTS" | grep -qi "^sitemap:"; then
  echo "WARNING: robots.txt missing Sitemap directive"
fi

if echo "$ROBOTS" | grep -qiE "dissallow|disalow"; then
  echo "ERROR: misspelled directive - rule will be ignored"
  exit 1
fi

The first check is the one that matters. It costs nothing and it catches the single most expensive mistake on this page.

A robots.txt template for Next.js

A solid starting point for most applications:

import type { MetadataRoute } from 'next';

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      {
        userAgent: '*',
        allow: '/',
        disallow: ['/api/', '/dashboard/', '/auth/', '/admin/'],
      },
    ],
    sitemap: 'https://yourdomain.com/sitemap.xml',
  };
}

Generating the file this way is better than a static one: it's type-safe, version-controlled with your code, and can vary per environment - which is exactly where the testing above earns its keep.

The five-point checklist

Before every deploy, verify:

  1. No Disallow: / - unless you're intentionally blocking a staging environment
  2. CSS and JS are reachable - don't block /_next/, /static/, or /assets/
  3. Sitemap directive is present - and the sitemap it points at is valid
  4. Paths use trailing slashes - /api/ not /api
  5. Environment logic is correct - production doesn't inherit staging rules

The fastest way to check

If you want to skip the manual testing and validate your robots.txt in seconds, paste your URL into the LintPage Robots.txt Validator. It catches syntax errors, overly broad blocks, missing sitemaps, and conflicting rules automatically.

For the wider pre-deploy picture, the pre-launch SEO checklist covers robots.txt alongside the other checks that silently block indexing, and our own 47-day noindex incident is what happens when a crawl directive ships wrong and nobody notices.

§ try this tool
Robots.txt Checker
Validate your robots.txt file for syntax errors and blocking rules.
try it free →
§ about the author
Marius OrzaruFounder, LintPage (BludeskSoft)

I built LintPage after a single stray noindex tag slipped into production and quietly cost us 47 days of organic traffic. It now runs the 60 automated checks I wish we had run before that deploy.

LinkedIn →

Get notified when we publish new posts.

§ run all 60 checks at once

Want the full picture? Stop checking one thing at a time.

Get a complete pre-launch SEO audit of your site with a single click.

run a full audit →
lintpage

Pre-launch SEO linting for developers. Catch disasters before they ship.

Product

  • Overview
  • Pre-launch checks
  • Full audit

Free tools

  • Meta tag checker
  • Robots.txt validator
  • AI crawler checker
  • llms.txt checker
  • llms.txt generator
  • og checker
  • Sitemap validator
  • Heading checker
  • SSL checker
  • Redirect checker
  • Structured data validator
  • Broken link checker
  • Core Web Vitals checker
  • Security headers checker
  • Canonical tag checker
  • Favicon checker
  • All tools →

Resources

  • Search Console errors
  • llms.txt guides
  • All checks
  • Blog
  • About
  • RSS feed
  • Contact

Legal

  • Privacy
  • Terms
© 2026 lintpage. All rights reserved.built after one too many post-mortems.