Find files, editable templates and browser test targets by what you need to make or test. The directory below is cut by format; the two collections under it cut the same library by subject and by workflow.
A robots.txt that disallows the entire site for every crawler while allowing one media path for a single image agent. For testing full-block handling and the common misconception that a Disallow removes a URL from a search index.
A robots.txt carrying three different crawl-rate hints - an integer Crawl-delay, a fractional one alongside Request-rate and Visit-time, and a very large one - none of which are part of RFC 9309. For testing that a crawler reads or ignores rate hints without dropping the Disallow rules that share the group.
A robots.txt using upper-case, mixed-case and indented field names, a tab-indented rule, a value with no space after the colon, and both trailing and full-line comments. For testing that field names are treated case-insensitively while path values stay case-sensitive.
A robots.txt declaring five sitemaps - before the first group, inside two different groups, in lower case, gzipped, and on another host. For testing that a discovery crawler collects Sitemap as a file-global field instead of scoping it to the group it sits in.
A robots.txt where Googlebot is named by two separate groups and every Disallow has an equal-length Allow competing with it. For testing group merging, case-insensitive product tokens, and the rule that the most specific match wins with Allow breaking ties.
A robots.txt in which real Disallow rules are surrounded by Noindex, Host, Clean-param, Nofollow and two invented fields, several with inline comments. For testing that a parser skips fields it does not implement instead of aborting or mis-binding the rules that follow.
A robots.txt that begins with a UTF-8 byte-order mark, uses CRLF line endings, and leaves trailing spaces after two rule values. For testing the byte-level tolerances RFC 9309 requires - a leading BOM must be discarded rather than glued onto the first field name.
A robots.txt built entirely from pattern rules: `*` inside a path, `$` anchoring the end of a URL, a query-string pattern, and an Allow that carves one file out of a disallowed subtree. For testing that a robots parser implements RFC 9309 path matching rather than prefix comparison.
The single most common broken robots.txt in the wild: a server that answers /robots.txt with its HTML 404 page instead of a 404 status. For testing that a crawler treats unparseable markup as 'no robots.txt' and allows the site rather than inventing rules from tag names.
A deliberately invalid robots.txt: a rule before any group, an empty product token, an absolute URL where a path belongs, whitespace before the colon, a non-numeric Crawl-delay, and a line with no colon at all. For testing that a parser degrades line-by-line instead of failing the whole file.