Which AI crawlers may read your site, decided and enforced
AI Search Optimization (GEO / AEO) · llms.txt and AI crawler access setup
The split between assistants that send traffic back and ones that only take it, written into robots.txt and enforced at the edge.
About this service
Allow OAI-SearchBot, PerplexityBot and Claude-SearchBot. Block GPTBot, CCBot and Bytespider. That split is right for most companies I open, because the first group fetches a page in order to cite it with a link back and the second fetches it to absorb it. The engagement exists because "right for most companies" is not something you can take to a board, and because a large share of the sites I look at are enforcing the opposite of what they believe they enforce.
The usual cause is not robots.txt. It is the CDN. Cloudflare began blocking AI crawlers by default for new domains in 2025, and many teams inherited that setting without anyone deciding it. A robots.txt can say Allow while the edge returns 403 to the same user agent, and nobody inside the company knows which is true until someone reads the server logs. That is the first thing I do.
About llms.txt:
No major assistant reads it as an ingestion source. Not OpenAI, not Google, not Anthropic, not Perplexity. If you were told otherwise, you were sold something. I still ship one in about half of these engagements, for reasons that have nothing to do with those four: internal teams use it as the one honest index of what the site actually contains, and some smaller retrieval tools do fetch it. If you want llms.txt because of a thread claiming it lifts ChatGPT visibility, say so before you pay and I will talk you out of it.
Who does this:
Me. I have run this decision inside publishers with licensing revenue to protect and inside SaaS companies whose pipeline now starts with somebody asking an assistant a question. Those two cases end in opposite policies, which is the point of deciding rather than copying.
The question the work answers:
For each named crawler, does letting it read this site make money or cost money, and which executive owns that answer. Not "is AI good for us".
How it is decided:
Ninety days of edge and origin logs, with every request claiming an AI user agent checked against the IP ranges the operators publish, OpenAI's gptbot.json and the equivalents from Anthropic, Perplexity and Common Crawl. Spoofed traffic is separated from real traffic first, because a policy written on unverified user agents is a policy written on noise. Then the referral side: assistants that cite send identifiable sessions, and those sessions are valued against the content they consumed. A crawler taking bandwidth and returning nothing measurable is a different decision from one returning sessions at a cost per session you can put next to paid.
What ships:
A decision memo, two pages, one line per crawler with an owner's name on it. robots.txt, written and deployed. Matching rules in whatever sits at your edge, Cloudflare WAF custom rules, Fastly VCL or Akamai, because robots.txt is a request and the edge is enforcement. A decision on nosnippet, max-snippet and data-nosnippet, since Googlebot cannot be separated from AI Overviews without also leaving Google's index, and that trade has to be made on purpose. A saved log query so the policy can be audited monthly by someone who is not me.
When the answer is inconvenient:
Sometimes it is: block Perplexity, accept the loss of referral sessions you can count, to protect a licensing position you cannot yet count. I write that down with both numbers attached and present it to whoever owns revenue, not to the search team. Where legal wants a blanket block and growth wants blanket access, I will not average them into a compromise nobody defends. Each side signs the line naming what it gave up.
Not included:
Licensing negotiation, legal drafting, bot-management vendor selection, and any promise that an assistant will cite you. Access is a precondition for citation. It is not a cause of it.
Turn this down if:
You want a robots.txt file emailed to you. That is a real need and this is the wrong way to buy it. You want every AI crawler blocked and are not open to evidence that it costs you. Or the person who would have to sign the memo has not agreed to read it, in which case buy nothing until they have.
Scope
- Target market
- Worldwide, United States, United Kingdom, Israel
- Working language
- English, Hebrew
- Industry
- B2B SaaS, Developer tools, Ecommerce and DTC, Marketplaces, Media and publishing
- Engagement model
- One-off project
- Turnaround
- 1 week, 2 weeks
- Seller type
- Freelancer, Fractional executive
What the seller needs from you
- 1Which domains and subdomains are in scope, and who signs the access policy for them?
- 2Can you give read access to 90 days of edge and origin logs?
- 3What sits in front of the site: CDN, WAF, bot management?
- 4Is any content licensing conversation open or planned?
- 5What do you currently believe your robots.txt allows?
Asked at checkout. Delivery time starts once you answer, not when you pay.
Reviews
No reviews on this service yet.
Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.
Starting at $8,500