By Greg Nowak. Updated 3 October 2026.
Your robots.txt file may be only a few lines long, but the decisions behind it are becoming harder. A business might welcome search traffic, want its expertise cited in AI answers, decline model training and negotiate terms for reuse of its original research. One domain-wide “allow AI” or “block AI” setting cannot express all of that.
A crawler licensing register gives those decisions a home. It is a maintained record of which content may be used for which purpose, how that choice is communicated, who approved it and how the team checks it. An agency can use the same record to keep client policy, SEO work and infrastructure changes aligned.
Start with the use and the content
Separate public marketing pages from assets with different value or obligations: paid reports, product data, downloadable media and customer material. Then decide what each group is meant to do. The table is a starting point for that conversation, not a preset policy.
| Use | Decision to make | Evidence to keep |
|---|---|---|
| Conventional search | Should these pages be discoverable in search results? | Indexability checks and search crawler rules |
| AI answers | Is visibility in generated answers useful for this content? | Relevant crawler settings and example answer checks |
| Model training | Is this use permitted, refused or subject to terms? | Training controls, published terms and approval |
| User-requested retrieval | Can an assistant fetch this page when someone asks? | Access rules, bot verification and request logs |
| Licensed reuse | Are attribution, payment or a contract required? | License reference, asset scope and contract owner |
Translate the decision into the right control
robots.txt tells compliant crawlers which URLs they may request. It does not secure a file or guarantee that a page stays out of search. Put private content behind authentication. For a public page that should be excluded from Google Search, use an appropriate indexing control and make sure Google can crawl the page to read it.
Crawler names also need careful handling. Google’s crawler documentation says Google-Extended controls specified Gemini training and grounding uses without affecting Google Search inclusion or ranking. It is a robots.txt control token, not a separate HTTP user agent you can find in access logs. OpenAI’s crawler documentation distinguishes OAI-SearchBot for ChatGPT search from GPTBot for content that may be used in training. ChatGPT-User covers certain user-initiated visits; OpenAI says robots.txt may not apply to those requests.
For example, a site that wants to decline GPTBot access and the uses controlled by Google-Extended could add these groups to its reviewed robots.txt file:
User-agent: Google-Extended
Disallow: /
User-agent: GPTBot
Disallow: /This example also declines the Google-Extended grounding use; it is not a training-only switch. It says nothing by itself about OAI-SearchBot or Googlebot. Check the live file before adding an explicit search-crawler group: Google explains that a specific group does not inherit the wildcard group’s rules. An apparently harmless Allow: / can therefore undo path restrictions you intended to keep. Test important public and restricted paths after every edit.
Keep licensing separate from crawl access
A crawler permission and a reuse license answer different questions. If you offer content under defined terms, the RSL 1.0 specification provides a machine-readable XML format and a way to advertise a site-level license in robots.txt:
License: https://example.com/ai-license.xmlThat URL must lead to an actual license covering the assets and uses you intend. A reference alone does not create a deal or ensure every crawler will honor it. Have the commercial owner and appropriate counsel review terms involving payment, attribution or restrictions. Keep signed agreements and exceptions in the register alongside the public signal.
Make the register usable on a busy Tuesday
A spreadsheet is enough to begin. Give each row a content group and intended use, then record the decision, relevant crawler token, robots.txt rule or license URL, any CDN or WAF control, the person who approved it, the evidence source and the next review date. Link to the actual configuration or document rather than pasting a policy statement nobody can verify.
Start by capturing what visitors receive at /robots.txt, plus current edge rules and a small sample of request logs. Compare those with the intended decisions. If an agency manages the site, name who can change the CMS, CDN and licensing terms; otherwise, a routine deployment can quietly reverse a policy agreed elsewhere.
For monitoring, distinguish a declared user-agent from a verified crawler. User-agent strings can be copied. Use published IP information or your provider’s bot verification where available, and look for requests to paths you meant to restrict. Cloudflare’s guidance makes the operational distinction clear: robots.txt communicates preferences, while AI Crawl Control can enforce access rules at the edge.
Review when the business or the crawlers change
Revisit the register when you launch a valuable content type, sign a reuse agreement, change CDN settings or see a provider revise its crawler documentation. The review can be short: confirm the decision still makes commercial sense, inspect the live rules, test representative URLs and record what changed. That turns a fragile text file into a decision the business can explain.
If your crawler settings currently sit between marketing, operations and legal ownership, Greg can help map the content, agree the permissions and turn them into a tested, maintainable register. Get in touch to discuss a crawler permissions review.
Related on GrN.dk
- ChatGPT Visibility Without Opening Every Door: robots.txt Is Only the Start
- AI crawler policy now has verbs: separate search, RAG, and training
- Bot traffic now exceeds human traffic—and crawler rules are no longer optional
Need help with this kind of work?
Discuss a crawler permissions review Get in touch with Greg.