How to Use the llms.txt Generator, Step by Step
Reviewed by the OnlineFree.app team · Updated
Key points
- The llms.txt Generator turns a site name, URL and allowed paths into an upload-ready llms.txt plus a matching robots.txt snippet.
- Allowed paths are capped at 50 lines, with blank lines and duplicates stripped automatically.
- Upload llms.txt to your domain root and serve it as text/plain at /llms.txt.
- Append the robots.txt snippet to your existing file instead of replacing it.
- llms.txt is a convention, not a guaranteed citation signal — verify with server logs.
What the llms.txt Generator produces
The llms.txt Generator turns three plain facts into a valid llms.txt file and a matching robots.txt snippet: your site name, your site URL, and the section paths you want answer engines to read. Three fields, one Generate button, roughly a minute of work, and no account.
The file it outputs follows the llms.txt convention — an H1 carrying your site name, a one-line summary starting with a > blockquote marker, an '## Allowed paths' bullet list of absolute URLs, and a short '## Notes' line stating the file exists for AI and answer-engine crawlers. You can copy it from the code block or download it as llms.txt in text/plain.
The second tab covers the crawler half: User-agent blocks that allow GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot with 'Allow: /', a 'Sitemap:' line derived from your site URL, and a comment naming your llms.txt. The underlying convention is documented at llmstxt.org.
How do you generate an llms.txt file step by step?
Start with the site name exactly as you want it to appear in the H1 — 'Acme Analytics', not 'acme-analytics-home'. That string becomes the heading and feeds the summary line, and it is the first thing a reviewer sees when they open your file during a GEO or AEO audit.
Next, paste the site URL, for example https://acme.com. The generator normalizes it to scheme plus host and strips the trailing slash, so every link in the file is a clean absolute URL and the Sitemap line stays consistent. Enter the exact host you publish on: if the site resolves at www.acme.com, use that.
Then list allowed paths, one per line — '/docs', '/pricing', '/about', or full URLs if you prefer. Blank lines and duplicates are removed, bare paths get your site URL prefixed, and the list is capped at 50 entries. Press Load example to see the shape first, hit Generate, then copy or download the result.
Where do you upload llms.txt and robots.txt?
Put llms.txt at your domain root so it is reachable at https://acme.com/llms.txt, served as text/plain. Root placement matters: crawlers and audit tools look at the fixed path /llms.txt instead of searching for it, and each subdomain is a separate host with its own file.
For robots.txt, append the generated blocks to whatever is already published. Overwriting the file can silently drop rules you depend on, duplicate 'Sitemap:' lines trip up some parsers, and retyping user-agent groups by hand is how typos get in. The file's syntax rules are standardized in RFC 9309.
Both outputs come from the same single-screen tool, like the rest of the free online tools on OnlineFree.app — no install, no signup. If you keep staging copies or a docs subdomain and need filenames tidied before uploading a batch, the Bulk File Renamer handles that side of the job.
Mistakes that make an llms.txt file useless
The most common mistake is listing paths you don't actually want crawled. 'Allowed paths' is a signpost, not access control: adding /admin or a private beta URL advertises it. Keep the list to pages that are public, canonical and genuinely useful to a machine reader.
Second, host mismatches. If your canonical host is www.acme.com but you generate with acme.com, every absolute URL in the file points at a redirecting host. Third, listing URLs that redirect or that carry tracking query strings, which multiply near-duplicates without adding meaning.
Fourth, treating 50 lines as a target. Three to eight well-chosen paths — docs, pricing, about, one key landing page — stay accurate far longer than a 50-line dump nobody updates. Regenerate the file whenever your site structure changes, the same way you would update a sitemap.
What llms.txt can't do for you
llms.txt is a convention, not a standard enforced by any search engine or AI vendor. As of 2026 there is no public commitment from OpenAI, Anthropic, Google or Perplexity that their crawlers read it, and no tool can promise that publishing one will get your site cited in an AI answer. What it removes is ambiguity.
Robots directives only bind crawlers that choose to honor them. A bot that ignores robots.txt will ignore the snippet too, and an adversarial scraper was never going to ask permission. Treat both files as clarity for well-behaved crawlers and as documentation for the humans running GEO and AEO audits.
Verify rather than assume: fetch your live /llms.txt and /robots.txt in a browser, confirm each returns 200 with a text/plain content type, then watch server logs over a few weeks for GPTBot, ClaudeBot or PerplexityBot hits. Silence doesn't mean harm, but don't report a win you can't measure.
Frequently asked questions
What does the llms.txt Generator actually create?
It creates two copy-pasteable blocks: a complete llms.txt file (H1 site name, a one-line blockquote summary, an '## Allowed paths' list of absolute URLs, and a short notes line) and a robots.txt snippet that allows GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot while pointing to your sitemap and llms.txt. Both are ready to upload.
Is the llms.txt Generator free, and do I need an account?
It is free and requires no signup. You supply the site name, site URL and allowed paths, and the page normalizes your URL and builds the output before you copy or download it. There is no dashboard and no email capture; the only structural limit is the 50-line cap on the allowed-paths field.
Do I replace my existing robots.txt with the snippet?
No. Append the generated user-agent blocks to your existing robots.txt rather than overwriting it, because deleting rules you already rely on can expose directories you meant to block. Keep a single 'Sitemap:' line, watch for duplicate user-agent groups, and confirm the merged file still follows RFC 9309.
How many paths can I list in the llms.txt Generator?
The allowed-paths field accepts up to 50 non-empty lines. Blank lines and duplicate entries are dropped automatically, and each line becomes an absolute URL, so '/docs' turns into 'https://acme.com/docs' while full URLs you paste are kept as-is. Fewer, well-chosen paths usually read better than an exhaustive list.
Will adding llms.txt get my site cited by ChatGPT or Perplexity?
No tool can promise that. llms.txt is a convention rather than a ranking factor, and as of 2026 there is no public commitment from the major answer engines to read it. What it does is remove ambiguity: crawlers get a short machine-readable map of your best pages plus a clear robots.txt posture, which is what GEO and AEO audits check.