How to Use the Robots.txt & LLMs.txt Generator, Step by Step
Reviewed by the OnlineFree.app team · Updated
Key points
- The Robots.txt & LLMs.txt Generator produces both files from one screen with no signup, crawling or build step.
- Its default AI policy blocks training crawlers such as GPTBot and CCBot while still allowing OAI-SearchBot and PerplexityBot.
- robots.txt is a voluntary convention, so blocked crawlers may still access pages; use authentication for private data.
- The generated llms.txt is a starter file based on a proposal, not a formal standard or a promise of AI citation.
- Verify both files at your domain root after upload and regenerate them whenever your URL structure changes.
What does the Robots.txt & LLMs.txt Generator do?
It converts a domain, a short list of disallowed paths and one AI-crawler policy choice into two ready-to-paste files: a complete robots.txt and a matching starter llms.txt. Everything updates live on every keystroke in a single screen, with no signup, no crawling of your site and no build step.
The two outputs sit in monospace panels with a filename header, a copy button and a download button each. The robots.txt panel contains User-agent groups, your Allow and Disallow rules for the * wildcard and for each selected AI crawler, plus a Sitemap line derived from the origin you typed.
The llms.txt panel is Markdown: an H1 built from the domain, a > summary line, a "Do not crawl" section mirroring your disallowed paths, and a "## Pages" section with placeholder links to fill in. Treat it as a starter file you edit, not a finished document.
You can open it at Robots.txt & LLMs.txt Generator; it is one of the browser-based tools on OnlineFree.app and requires no account.
Step 1: Enter your site origin
Paste either example.com or https://example.com/ into the first field and the tool normalises the protocol and trailing slash for you. Only the origin is needed — no paths, no query strings — because that single value feeds the Sitemap line in robots.txt and the header of llms.txt.
Subdomains matter here. blog.example.com produces different output than example.com, and each host is a separate crawler namespace with its own robots.txt. Staging hosts work the same way, so don't copy production rules onto staging.example.com without checking the paths still apply.
Validation errors appear directly under the offending field, so an empty or malformed origin is flagged before you copy anything into your repo.
Step 2: List the paths you want blocked
Type one path per line in the textarea — /admin/, /cart, /private/, *.json$ are typical entries. Blank lines and lines starting with # are ignored, so you can group rules into commented blocks. You can also add a short trailing comment after a path to record why it is blocked, which helps when a rule outlives the person who wrote it.
Wildcards pass straight through to the output, but support varies: some crawlers interpret * and $ as patterns while others treat those characters literally. Don't rely on pattern matching alone for anything that genuinely matters.
It also helps to remember what Disallow actually does. It asks crawlers not to fetch a URL, which means a blocked page can still appear in results if other pages link to it — and a noindex tag can't be read on a page crawlers are forbidden to fetch. For genuinely private data, use authentication instead.
Step 3: Pick an AI crawler policy
The dropdown has three options and defaults to block_training_allow_search. Choosing allow_all adds no AI-specific rules, leaving every crawler governed only by your * group. block_training_allow_search writes Disallow rules for GPTBot, Google-Extended, CCBot, anthropic-ai and Bytespider while leaving OAI-SearchBot and PerplexityBot allowed. block_all disallows every known AI crawler the tool lists.
The distinction matters because "AI crawler" covers two different jobs. Training crawlers collect text for model training, while search and retrieval bots fetch a page to answer a specific query and cite the source. Blocking the first group while allowing the second is the common middle ground for publishers who want citations but not training use.
None of these rules are enforceable on their own. robots.txt is a voluntary protocol described in RFC 9309, and crawlers that ignore it — including some scrapers — are unaffected by anything you write. Treat the policy choice as a signal rather than a lock.
Step 4: Copy, download and deploy both files
Use the copy or download button in each panel. robots.txt must be served from the domain root at https://example.com/robots.txt; a copy buried in a subdirectory is ignored. Serve it as plain text with a 200 status, and check that the URL doesn't redirect — a redirect or a typo in the filename is a common reason rules seem to have no effect.
llms.txt follows the same placement convention at /llms.txt on the root. It comes from the llms.txt proposal, which is a convention adopted by early adopters rather than a formal standard as of 2026, so expect the format and how it is consumed to keep changing.
After uploading, open both URLs in a browser and confirm you see exactly the text you generated. Re-run the generator whenever your site structure changes, since paths blocked a year ago may no longer match. If you keep screenshots of the deployed files for a changelog or ticket, Free Online Photo Resize Tool can shrink them first.
What are the most common robots.txt mistakes?
Over-blocking is the big one. Disallowing /assets/, /css/ or entire language folders stops crawlers from fetching the files needed to render a page, which can make a perfectly good page look broken to search engines.
The second mistake is treating the file as a privacy control. It isn't one. Search engines that obey the file still know the URL exists, and a disallowed URL can be indexed from inbound links. Logins, tokens and internal documents belong behind authentication or off the public host entirely.
Third, the output is only as good as the input. This generator does not crawl your site, check whether a path exists, confirm that a given crawler honours a rule, or guarantee that any AI system will read your llms.txt. It removes the syntax memorisation; reviewing the result against your real URL structure is still your job.
Frequently asked questions
Is the Robots.txt & LLMs.txt Generator free, and do I need an account?
Yes. The Robots.txt & LLMs.txt Generator is free and needs no account or email address. It runs in your browser, does not crawl your site and does not store your domain or paths, so there is nothing to sign up for. You can copy or download both files immediately after typing.
What is the difference between block_training_allow_search and block_all?
block_training_allow_search writes Disallow rules for GPTBot, Google-Extended, CCBot, anthropic-ai and Bytespider but keeps OAI-SearchBot and PerplexityBot allowed, so search and retrieval crawlers can still fetch pages. block_all disallows every known AI crawler in the tool's list. allow_all adds no AI-specific rules at all, leaving only your * group in charge.
Does blocking AI crawlers in robots.txt stop AI systems from using my content?
Not reliably. robots.txt is a voluntary protocol under RFC 9309, and only cooperating crawlers honour it. Content already collected, licensed, or scraped by crawlers that ignore robots.txt is unaffected by new rules. Use robots.txt to express a preference, and rely on authentication or licensing terms for content you must protect.
Where should robots.txt and llms.txt be placed on my domain?
robots.txt must sit at the domain root, for example https://example.com/robots.txt, and return a 200 status as plain text; copies in subdirectories are ignored. llms.txt is conventionally placed at /llms.txt on the same root. Both are host-specific, so example.com and blog.example.com each need their own files.
Can I use wildcards and comments in the disallow paths field?
Yes. Enter one path per line, and lines starting with # or left blank are ignored. You can append a short comment after a path to document the reason it is blocked. Wildcards such as *.json$ pass through to the output, but crawler support varies, so avoid depending on patterns for critical rules.