← Robots.txt & LLMs.txt Generator

Robots.txt & LLMs.txt Generator: Tips and Common Mistakes

Reviewed by the OnlineFree.app team · Updated

Key points

  • The Robots.txt & LLMs.txt Generator turns a domain, disallow paths and one AI policy into two paste-ready files.
  • Its default AI policy blocks training crawlers like GPTBot and CCBot but keeps OAI-SearchBot and PerplexityBot allowed.
  • Robots.txt controls crawling, not indexing — a blocked URL can still appear in search results via external links.
  • llms.txt is a proposal as of 2026, while robots.txt is standardized in RFC 9309 and honored by mainstream crawlers.
  • Always upload robots.txt to the site root and verify the Sitemap line and rules before relying on them.

What does the Robots.txt & LLMs.txt Generator output?

The Robots.txt & LLMs.txt Generator turns one site origin, a list of disallow paths and a single AI-crawler policy choice into two ready-to-paste files: a complete robots.txt and a matching llms.txt starter in Markdown. Generation happens in the browser on one screen — no signup, no crawl of your live site, no build step.

The robots.txt half is the enforceable part: User-agent groups with Allow and Disallow rules for * plus each AI crawler your policy selects, and a Sitemap line derived from the origin you typed. The llms.txt half is documentation: an H1 built from the domain, a > summary line, a 'Do not crawl' section mirroring your disallow paths, and a '## Pages' section with placeholder links you replace with real URLs.

Because the preview updates on every keystroke, you can watch what adding a path or flipping the AI policy does before you copy anything. That matters most when you're writing rules for GPTBot, ClaudeBot, CCBot or PerplexityBot and would rather not guess at exact user-agent spellings.

Your first 60 seconds: domain, paths, policy

Start with the origin — example.com and https://example.com/ both work, because protocol and trailing slash are normalized automatically. That one value feeds the Sitemap line and the llms.txt header, so if your sitemap lives at a non-standard filename, adjust it once after copying.

Then add disallow paths, one per line. Blank lines and lines beginning with # are ignored, and an optional trailing comment after a path lets you record why it exists — useful six months later when nobody remembers why /cart was blocked. Precision beats volume here: Disallow: /admin also blocks /admin/, /administrator and /admin-panel, while Disallow: /admin/ covers only the directory and everything beneath it.

Wildcards work in mainstream crawlers. Entering *.json$ in the path box stops crawlers fetching raw JSON endpoints, because * matches any string and $ anchors the match to the end of the URL. RFC 9309 codifies both operators, so you aren't relying on undocumented behavior — though a handful of minor bots still only understand plain prefixes.

Which AI-crawler policy should you choose?

There are three options and one meaningful difference between them. allow_all writes no AI-specific rules, so ordinary search crawlers and AI agents get identical treatment. block_training_allow_search — the default — disallows GPTBot, Google-Extended, CCBot, anthropic-ai and Bytespider while leaving OAI-SearchBot and PerplexityBot allowed. block_all disallows every known AI crawler in the tool's list.

The default suits most publishers: you opt out of bulk model-training crawling while keeping the crawlers that can still surface and cite your pages in answer engines. If your revenue depends on being quoted by AI assistants, block_all removes a distribution channel; if your content is licensed or paywalled, the opposite trade-off may be correct.

Two honest caveats. Robots.txt is a voluntary protocol — polite crawlers honor it and impolite ones ignore it entirely, so it is a signal, not a lock. And vendors rename user-agents over time, so treat any generated list, including ours, as a snapshot and re-verify it periodically; as of 2026 the AI bot names are still shifting.

Common robots.txt mistakes to avoid

The mistake we see most is broken grouping. Every rule line beneath a User-agent line belongs to that agent until the next User-agent line appears, so a missing User-agent row silently attaches your rules to the wrong bot, and one group listing several agents shares a single rule set — exactly how the generator emits a group per policy tier.

The second classic is blocking assets. Disallowing /assets/, /wp-content/uploads/ or *.css stops crawlers fetching the files that make your pages render, which hurts how search engines evaluate them. Keep crawl rules for pages, directories and parameters; leave the CSS, JS and images alone.

Then there's the crawling-versus-indexing confusion. A Disallow line prevents a crawl of a URL; it does not guarantee the URL leaves an index, because a blocked page can still be listed if other sites link to it. Use noindex, canonicals and removal requests for indexing control — and remember noindex only works on pages crawlers are allowed to fetch.

Does llms.txt actually do anything yet?

Not in the enforceable sense. As of 2026, llms.txt is a proposal rather than a standard: it is a Markdown convention published at llmstxt.org asking site owners to describe their site for AI systems, and no major crawler is currently obligated to fetch or obey it. Robots.txt, by contrast, is standardized in RFC 9309 and honored by mainstream search crawlers.

That does not make the file pointless. The generated llms.txt is a compact, human-readable brief — an H1 from your domain, a > summary line, a 'Do not crawl' list mirroring your disallow paths, and a '## Pages' section where you paste the handful of URLs that best represent the site. Many teams publish it as cheap, structured guidance and as a statement of intent.

So use the two files for different jobs. The llms.txt earns its keep as documentation for people and for agents that choose to read it; the robots.txt earns its keep by actually restricting crawler access. Ship both, but don't expect the Markdown file to enforce anything on its own.

Verify before you ship

Upload robots.txt to the root of the exact origin — https://example.com/robots.txt — because crawlers do not look anywhere else, and each subdomain needs its own copy. After deploying, test a few representative URLs using the tools described in Google Search Central's robots.txt documentation, and confirm the Sitemap line matches the sitemap that actually exists.

Keep staging separate from production. We generate one file per environment and name them robots.staging.txt, robots.preview.txt and so on; if you batch those alongside other per-environment files, a Bulk File Renamer handles the renaming faster than doing it by hand.

Finally, don't let any generator, including ours, have the last word. Diff the output against your previous file, keep the result in version control, and re-check whenever a crawler you care about announces a new user-agent name. The rest of the free, browser-based OnlineFree.app toolset works the same way — quick output, with you doing the final verification.

Frequently asked questions

Is the Robots.txt & LLMs.txt Generator free, and do I need an account?

It is free and requires no account. You type a domain, optional disallow paths and an AI-crawler policy, and the robots.txt and llms.txt files are assembled in the browser as you type — no crawling of your site, no build step, no upload. Both outputs have copy and download buttons.

What is the difference between robots.txt and llms.txt?

Robots.txt is a crawler-access file standardized in RFC 9309 and honored by mainstream search crawlers, so it can actually restrict access. llms.txt is a Markdown proposal published at llmstxt.org that describes a site for AI systems; as of 2026 no major crawler is obligated to read or follow it, so it works as guidance rather than enforcement.

Will blocking GPTBot keep my site out of AI answers?

Not automatically. The default policy in the Robots.txt & LLMs.txt Generator blocks training-oriented crawlers such as GPTBot, Google-Extended, CCBot, anthropic-ai and Bytespider, but leaves OAI-SearchBot and PerplexityBot allowed, which are the agents tied to surfacing and citing pages. Choosing block_all disallows every listed AI crawler and may reduce your AI visibility.

Do I need a robots.txt file at all?

No — crawlers assume everything is allowed when no robots.txt exists. It is worth adding anyway to point to your sitemap and to exclude paths that waste crawl budget or leak data, such as /admin/, /cart, internal search results and raw JSON endpoints. A short, accurate file beats a long speculative one.

Can robots.txt rules use wildcards like * and $?

Yes. RFC 9309 defines both operators, and mainstream crawlers support them: * matches any sequence of characters and $ anchors a pattern to the end of the URL, so a rule like *.json$ blocks raw data endpoints. A few minor bots only understand plain path prefixes, which is one reason to test important rules after publishing.

References

Try Robots.txt & LLMs.txt Generator free — no sign-up, works in your browser
Open the tool →

More free tools

Step-by-step guides in our blog & guides.

Add Numbers Online Calculator Css Gradient Generator Online Wie Viel Prozent Sind Rechner Percentage Calculator Free Online Lease Agreement Template Patreon Revenue Calculator Free Online Photo Resize Tool محول صيغ الفيديو اون لاين Compress Image Size Online Free To 100kb Free Programs To Resize Images