# LLM SEO: What It Is, What's Different, and How to Do It (2026)

> LLM SEO means getting your brand cited by ChatGPT, Claude, and Perplexity. What changes vs traditional SEO, what Google says to skip, and how to do it.

- Published: 2026-08-15
- Updated: 2026-08-15
- Author: Samy BEN SADOK
- Canonical: https://geotoolbox.ai/blog/llm-seo

---

LLM SEO is one of the most oversold terms in marketing right now, and one of the most misunderstood. Strip out the hype and a straightforward discipline sits underneath: getting your brand named when ChatGPT, Claude, Gemini, and Perplexity answer a question. This guide separates what works from what vendors are selling, including what Google itself says you can safely ignore.

## What LLM SEO Is (and the Confusion to Clear First)

LLM SEO is the practice of structuring your content and brand presence so large language models like ChatGPT, Claude, Gemini, and Perplexity mention and cite you when they answer a question. The goal is not a blue-link ranking. It is being the source the model pulls from.

Before anything else, clear up the phrase, because it is ambiguous in a way that sends people down the wrong path. Two people search "LLM SEO" and mean opposite things:

- **Optimizing to appear in LLMs.** You want ChatGPT to name your brand when someone asks it for a recommendation. That is what this article is about.
- **Using an LLM to do your SEO.** You want ChatGPT to write meta descriptions or draft content faster. That is a writing-workflow question, not a visibility one.

So, can ChatGPT do SEO? It can help you produce and audit content, but "LLM SEO" as a discipline means the first thing: earning a place inside the answer. Keep the two apart and the rest of the field stops sounding like noise.

That noise includes a pile of acronyms. They all describe the same underlying job, getting your brand surfaced and cited inside AI-generated answers, from different angles.

<table>
<thead>
<tr><th>Term</th><th>What it emphasizes</th></tr>
</thead>
<tbody>
<tr><td><strong>LLM SEO</strong></td><td>The models specifically (ChatGPT, Claude, Gemini)</td></tr>
<tr><td><strong>LLMO</strong> (LLM optimization)</td><td>Same as LLM SEO, phrased around "optimization"</td></tr>
<tr><td><strong>GEO</strong> (generative engine optimization)</td><td>Generative answer engines broadly</td></tr>
<tr><td><strong>AEO</strong> (answer engine optimization)</td><td>Direct-answer surfaces like AI Overviews and featured snippets</td></tr>
<tr><td><strong>AI search optimization</strong></td><td>The plain-English umbrella</td></tr>
</tbody>
</table>

Vendors will sell you separate budgets for each. Do not buy it. We map the full vocabulary in [what LLMO is](https://geotoolbox.ai/blog/what-is-llmo) and the [GEO vs AEO vs SEO breakdown](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo); pick one word for your deck and move on. What matters is the work, and it starts with what changed under the hood.

## What Changes When a Language Model Ranks You: Training vs Retrieval

Here is the one idea most LLM SEO guides blur, and getting it straight resolves half the arguments in the field. A language model can surface your brand through two completely separate paths, and they run on different clocks.

<figure>
  ![LLM SEO's two paths: training, which moves in model generations, and retrieval, which moves in days.](/blog/llm-seo/llm-training-vs-retrieval-paths.png)
  <figcaption className="mt-3 text-center text-sm text-gray-500">The two ways a language model can surface your brand, and why they move at completely different speeds.</figcaption>
</figure>

**The training path.** During pretraining, a model reads a large slice of the public web and absorbs patterns about who is associated with what. If enough sources describe you as a leading option in your category, that association becomes part of the model's baseline knowledge. You cannot edit this directly, and it moves at the speed of model releases: months, sometimes longer. This is why a brand that was quiet a year ago can still be missing from a model's "from memory" answer today.

**The retrieval path.** When a model runs a live search to answer you, a separate system fetches current pages, and the model writes its answer from what it just pulled. This is [retrieval-augmented generation](https://geotoolbox.ai/blog/what-is-rag), and it moves in hours to days, depending on how often the retrieval crawler visits you. Publish a strong, reachable page and it can be cited in a live answer long before any model is retrained. Our breakdown of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) walks through the full loop.

Once you see the two paths, the common confusions dissolve. "I got cited in ChatGPT within a day" and "AI visibility takes months to build" are both true: the first is retrieval, the second is training. Freshness matters most on the retrieval path, especially for time-sensitive queries, and barely at all on the training path.

Two more shifts follow from this. There is no ranked list of ten blue links to climb; there is an answer, and you are either named in it or you are not. And the win is a citation, not a click, which changes how you prove the work is paying off (more on that below).

## LLM SEO vs Traditional SEO (and GEO, AEO)

Start with the uncomfortable part for anyone selling LLM SEO as a brand-new discipline: most of it is traditional SEO. If a crawler cannot fetch and render your page, no retrieval layer can pull it into an answer. Crawlability, fast rendering, server-side HTML, and clean information architecture are the entry ticket, not a bonus.

Google says this plainly. Its own guidance states that "the best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems," per [Google Search Central's AI-features guidance](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide). For Google's AI Overviews and AI Mode specifically, LLM SEO is SEO.

What changes is the last stretch: how the answer is assembled and what gets you named in it.

<table>
<thead>
<tr><th>Dimension</th><th>Traditional SEO</th><th>LLM SEO</th></tr>
</thead>
<tbody>
<tr><td><strong>What you optimize</strong></td><td>A page, for a keyword</td><td>A page and your off-site footprint, for a question</td></tr>
<tr><td><strong>The win</strong></td><td>A ranked blue link</td><td>A citation or mention inside the answer</td></tr>
<tr><td><strong>Who decides</strong></td><td>One ranking system (Google)</td><td>Several engines with different indexes and rules</td></tr>
<tr><td><strong>Authority signal</strong></td><td>Backlinks</td><td>Backlinks plus unlinked mentions and third-party consensus</td></tr>
<tr><td><strong>How you measure</strong></td><td>Rankings, clicks, impressions</td><td>Mention rate, citation rate, share of voice</td></tr>
<tr><td><strong>Timeline</strong></td><td>Weeks to months</td><td>Days (retrieval) to model generations (training)</td></tr>
</tbody>
</table>

Most of the difference sits in two rows. First, different engines do not reprint Google's top ten. They use their own indexes and source-selection rules, so ranking first on Google helps but does not guarantee you show up in ChatGPT (which draws on a mix of Bing and its own index) or Perplexity (which runs its own crawler). Second, authority stops being only about links to your site. A model builds its picture of you from what the rest of the web says, which is why the work leaks off your own domain (more on that in the playbook).

If you are deciding where to spend, this is a rebalancing of one budget, not a second one. We lay out [how to split your GEO budget](https://geotoolbox.ai/blog/geo-vs-seo), and cover the definitions in [what generative engine optimization is](https://geotoolbox.ai/blog/what-is-geo) and [what answer engine optimization is](https://geotoolbox.ai/blog/what-is-answer-engine-optimization). The tactics below are where the new part lives.

## The LLM SEO Playbook: Five Levers That Move Citations

Five levers do most of the work. None of them is a trick. The order matters, because the first one gates all the others.

<table>
<thead>
<tr><th>Lever</th><th>Why it works with an LLM</th><th>Go deeper</th></tr>
</thead>
<tbody>
<tr><td><strong>1. Reachability</strong></td><td>If AI crawlers can't fetch the page, the retrieval path never sees it</td><td><a href="https://geotoolbox.ai/blog/ai-crawlers">AI crawlers</a></td></tr>
<tr><td><strong>2. Answer-first structure</strong></td><td>A direct answer under a clear heading is a clean passage to lift into a response</td><td><a href="https://geotoolbox.ai/blog/content-chunking">Structuring for extraction</a></td></tr>
<tr><td><strong>3. Citable substance</strong></td><td>Statistics, citations, and quotes give a model concrete, attributable material</td><td><a href="https://geotoolbox.ai/blog/eeat-ai-search">E-E-A-T for AI</a></td></tr>
<tr><td><strong>4. Off-site presence</strong></td><td>Models build your entity picture from what other sites say about you</td><td><a href="https://geotoolbox.ai/blog/entity-seo">Entity SEO</a></td></tr>
<tr><td><strong>5. Freshness</strong></td><td>Retrieval leans on current pages for time-sensitive questions; evergreen authority ages more slowly</td><td><a href="https://geotoolbox.ai/blog/ai-content-optimization">AI content optimization</a></td></tr>
</tbody>
</table>

**Citable substance is the one with a controlled study behind it.** A widely cited [academic GEO study](https://arxiv.org/abs/2311.09735) tested content changes against a benchmark of real search queries and found that adding citations, quotations from credible sources, and statistics were the top-performing methods, with visibility gains of roughly 30 to 40 percent on its main metric. Plain-language, readable phrasing helped too. Keyword stuffing, the old SEO reflex, did not. So the highest-impact edit is often the least glamorous: put a real number, a named source, or a direct quote next to each claim.

**Answer-first structure needs a caveat**, because Google is explicit that you do not need to break content "into tiny pieces for AI to better understand it." That is true for Google's own AI features, which run on its core ranking systems. But for standalone retrieval engines like Perplexity and ChatGPT search, a clear answer immediately under a descriptive heading is simply a cleaner passage to extract. Write that way because it serves the reader, not because you are feeding a machine. The moment the structure hurts a human, you have gone too far.

**Off-site presence is the lever that breaks SEO habits.** In classic SEO, the unit of work is your page. In LLM SEO, a large share of the work is how you are represented on other people's pages: review directories, "best tools" listicles, comparison articles, YouTube videos and their transcripts, and community threads on Reddit and Quora. Models build a picture of your reputation from that chorus, and unlinked mentions can matter, not just links. This is closer to digital PR than to on-page optimization, and it is the part you control least.

Reddit deserves a specific warning, because "just post on Reddit" is the most-repeated tactic in this space and the fastest way to get burned. Communities do corroborate a brand, but manufacturing consensus with a throwaway account is easy to spot and gets the account banned. Participate honestly or not at all. And because AI search often [fans one question out into many](https://geotoolbox.ai/blog/query-fan-out), the goal off-site is to be present across the sub-questions a buyer actually asks, not to plant one link in one thread.

For the step-by-step version of all five, including exactly what to change on a page, follow our [playbook for optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search). The rest of this guide handles the parts that playbook assumes you have already sorted: whether the engines can reach you at all, whether any of this is working, and how you would know.

## Can AI Even Reach You? Reachability You Can Verify

This is the failure that wastes the most effort, because it is silent. You can nail every tactic above and still be invisible if the crawler that feeds the answer never reaches your page. Most guides list "check your robots.txt" and move on. The trap is more specific than that.

Different bots do different jobs, and blocking the wrong one has very different consequences. The major AI operators each run several, and the jobs differ (OpenAI documents its own set in its [crawler documentation](https://developers.openai.com/api/docs/bots)):

<table>
<thead>
<tr><th>User-agent</th><th>Operator</th><th>Job</th><th>Block it and...</th></tr>
</thead>
<tbody>
<tr><td><strong>GPTBot</strong></td><td>OpenAI</td><td>Gathers training data</td><td>You opt out of training data, not live answers</td></tr>
<tr><td><strong>OAI-SearchBot</strong></td><td>OpenAI</td><td>Indexes pages for ChatGPT search</td><td>Your pages stop being cited in ChatGPT search</td></tr>
<tr><td><strong>Google-Extended</strong></td><td>Google</td><td>Gemini training control</td><td>You opt out of Gemini training (not Google Search)</td></tr>
<tr><td><strong>PerplexityBot</strong></td><td>Perplexity</td><td>Indexes pages for Perplexity answers</td><td>Your pages stop being cited in Perplexity</td></tr>
<tr><td><strong>Claude-SearchBot</strong></td><td>Anthropic</td><td>Indexes pages for Claude search</td><td>Your pages stop being cited in Claude</td></tr>
</tbody>
</table>

One caveat that trips people up: the user-action fetchers (ChatGPT-User, Perplexity-User, Claude-User) grab a page when a person asks for it, and they do not all follow robots.txt. OpenAI says its rules may not apply to ChatGPT-User, and [Perplexity's user fetcher generally ignores robots.txt](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), while [Anthropic says Claude-User honors it](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler). So a robots.txt block is only fully reliable for the automated crawlers above. The common and costly mistake is blocking the retrieval crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) while thinking you only opted out of training. Your robots.txt "looks fine," and you have quietly stopped being cited.

The gate you did not set is worse. In July 2025, [Cloudflare began blocking AI crawlers by default for new domains](https://www.cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large/), calling itself the first major infrastructure provider to do so. Its rules have kept shifting since: under the [2026 category controls](https://blog.cloudflare.com/content-independence-day-ai-options/), from September 15, 2026 new domains block Training and Agent crawlers by default on ad-displaying pages while leaving Search crawlers allowed. The takeaway is not the exact default of the month; it is that if your site sits behind a CDN or WAF, an edge policy you never chose can decide which AI bots reach you. Cloudflare's own [June 2024 data](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/) found AI bots reached about 39 percent of the top million properties on its network, while just 2.98 percent of the top million had any AI-bot controls in place at all.

So verify, do not assume. Read your robots.txt line by line against the user-agents above, and our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) does exactly that: it shows which of 34 AI crawlers your robots.txt allows or blocks, with the exact line to fix. That covers the robots.txt layer only, though. To catch the WAF and CDN gates, grep your server logs for the fetching crawlers (GPTBot, OAI-SearchBot, PerplexityBot, Claude-SearchBot) to confirm they are getting a 200, and diff your raw HTML against the JavaScript-rendered version, since a crawler that does not run your JavaScript sees only the raw response. Across the sites we audit for AI visibility, a stale robots.txt rule or an edge default is a common and easily-missed reason a reachable-looking page never gets cited, and it is the cheapest thing to rule out first.

This reachability-first logic is also why we treat [llms.txt](https://geotoolbox.ai/blog/llms-txt) as a distraction for AI visibility. No major AI engine has confirmed using it as a citation signal, and an [Ahrefs study of 137,210 domains](https://ahrefs.com/blog/llmstxt-study/) found that 97 percent of the llms.txt files it saw received zero requests in May 2026. The file does have a use in some developer tooling, where coding agents can pull it to fetch docs, but that is not the same job as getting cited in an answer.

## Does LLM SEO Actually Work? An Honest Look at the Evidence

Yes, with a few caveats. The market is not in doubt: Alphabet's Sundar Pichai said Google's AI Overviews passed [2.5 billion monthly users](https://blog.google/innovation-and-ai/sundar-pichai-io-2026/) at Google I/O in May 2026, and AI Mode passed 1 billion. The question is not whether AI answers matter. It is whether your effort moves your slice of them.

The strongest number in the field is the GEO study's "up to 40%" visibility lift, which we cited above. Read what it actually measured: a gain on a research-built pipeline over the top Google results in 2023 and 2024, scored on a benchmark visibility metric, not commercial traffic to a business running on production ChatGPT or Perplexity. It tells you which content changes a model responds to. It does not promise you 40 percent more customers. Anyone quoting it as a business outcome is stretching it.

The most-repeated actionable claim in this space is that AI citations drop off sharply once content is more than about three months old. It gets restated across guide after guide, and we could not find a single named dataset behind it. It is also, conveniently, repeated most often by companies selling the tracking and refresh services it justifies. Treat it as a plausible hypothesis about the retrieval path, not a measured law.

The pattern to plan for in the near term is that your visibility goes up while your click traffic stays flat or dips. Being named in an answer is a branding and consideration win, and it often does not send a click at all. If you report AI SEO to a client or a boss using only sessions and clicks, a genuine win can look like a failure. The fix is to set the terms before you start: report citation share and branded-search lift alongside sessions, and baseline all three at the outset so the trade is visible rather than hidden.

So when is LLM SEO not worth it? If you have not fixed basic crawlability and indexing, or if you cannot commit to the off-site and content work over months, do the SEO fundamentals first and revisit. The harder question is whether your buyers use AI tools to research your category yet, and you can check it rather than guess: look at the share of traffic in GA4's AI Assistant channel (a lower bound, since app visits hide in Direct), spot-check a handful of buying-intent prompts in ChatGPT and Perplexity to see who gets named, and ask a few customers how they found you. LLM SEO amplifies a solid foundation. It does not substitute for one. For where the wider market sits, our [state of AI search report](https://geotoolbox.ai/blog/state-of-ai-search-2026) lays out the data.

## How to Measure LLM SEO (Without Fooling Yourself)

Measurement is where LLM SEO gets hard. The core problem is that these systems are non-deterministic: ask the same model the same question twice and you can get two different answers, with different sources cited. A single check tells you almost nothing, which is why our [guide to measuring AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) is built around repetition rather than one-off lookups.

So build a method instead of spot-checking:

1. Write a fixed, representative set of prompts your buyers ask, and freeze it. Changing the prompts changes the results, so the set has to stay stable to be a baseline.
2. Run each prompt several times per check, across each engine you care about, and record how often you are mentioned and how often you are cited with a link. One run is noise; the rate across runs is the signal.
3. Log which URL got cited, not just that you appeared. The page that gets pulled tells you what is working.
4. Hold a fixed competitor set and track your share of voice against it over time, so you are measuring movement, not a single snapshot.
5. Report per engine. ChatGPT, Gemini, and Perplexity draw from different indexes, so a blended number hides where you are winning and losing.

Keep three things separate while you do this. A **mention** is the model naming you in prose. A **citation** is a linked source. A **referral** is a click that actually lands on your site. They are not the same, and conflating them is how reports mislead.

That last one has a trap. The native ChatGPT, Claude, and Perplexity apps often do not pass a referrer, and GA4 files any referrer-less visit as **Direct**, so app-driven AI visits quietly land in Direct rather than showing up as AI referrals. GA4 added a built-in AI Assistant channel in 2026 that groups recognized engines automatically, which helps, but it is still referrer-based, so any AI visit without a referrer or a tag still falls to Direct. The one durable signal is a UTM in the link itself: ChatGPT appends `utm_source=chatgpt.com` to its citations, and a query-string tag survives referrer stripping. Treat any referral-based number as a floor, and pair it with prompt-based tracking that does not depend on a click happening at all. Our guides on [running an AI visibility audit](https://geotoolbox.ai/blog/ai-visibility-audit) and [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) go deeper.

One honest note, since we build a tool in this category. Every AI visibility tracker, ours included, samples a prompt set and estimates; none of us can see the real query volumes inside ChatGPT, and results shift run to run. A tracker's job is to make that sampling disciplined and repeatable, not to hand you a precise number that does not exist. Anyone promising exact AI query counts is selling certainty the systems do not expose.

## Common LLM SEO Mistakes

Most wasted effort in LLM SEO comes from a short list of avoidable errors:

- **Blocking the retrieval crawler by accident.** The quiet failure from earlier: you mean to opt out of training, but block a retrieval crawler too and drop out of live answers without noticing. Verify per user-agent.
- **Mass-producing AI-generated content.** The web is flooding with synthetic pages, which makes original data, first-hand testing, and named expertise more valuable, not less. Publishing more of the same slop moves you the wrong way.
- **Confusing AI-only markup with normal structured data.** Google is explicit that you do not need AI text files (its wording for files like llms.txt), AI-only markup, or extra schema.org markup to appear in its AI features. Standard structured data is a different thing: Google still recommends keeping it for rich results, and it stays ordinary SEO hygiene. Competitors pointing out that "nearly every cited page has schema" are describing a correlation, not proof that schema earns the citation. Keep your normal schema, skip the AI-specific inventions, and see [what schema does and doesn't do for AI search](https://geotoolbox.ai/blog/schema-markup-for-ai) for the detail.
- **Gating your best content behind a login and expecting citations.** If a crawler cannot read it, a model cannot cite it. The middle path is to expose a substantive, self-contained summary that can stand on its own and gate the rest, or use Google's documented markup for paywalled content. What you cannot do is serve AI bots the full text while blocking human readers; that is cloaking, and it is against search guidelines.
- **Chasing every engine at once.** Pick the two or three your buyers actually use and go deep. Spreading thin across eight surfaces wins none of them.
- **Naming yourself into a collision.** Inventing a branded category ("we're a marketing intelligence engine") can collide with an established term in the training data and confuse retrieval about what you even are. Describe yourself in words the model already associates with your category.

## Frequently Asked Questions

### Is SEO dead now that AI answers the question directly?

No. AI answers lean heavily on the same search indexes and live web, so crawlability, indexability, and quality content are still the foundation. What is dying is the assumption that a ranking equals traffic; a chunk of that attention now resolves inside the answer without a click. The work is shifting rather than ending.

### Can you get cited by AI without ranking well in Google?

Yes, and it happens often. Answer engines use their own indexes and can pull a well-structured, reachable page into an answer even if it sits on page two of Google. Strong Google rankings help your odds, but they are neither required nor a guarantee.

### Does ChatGPT cite the same pages that rank number one in Google?

Sometimes, but do not count on it. The overlap is partial and swings with the query: informational questions tend to pull more from the familiar authorities, while commercial and comparison prompts often surface listicles, forums, and review pages that never ranked first. Ranking well is a positive signal, but plan for the AI surface as its own target rather than assuming your Google winners carry over.

### Do I need an llms.txt file for LLM SEO?

No. No major AI engine has confirmed using llms.txt as a citation signal, a large Ahrefs study found 97 percent of published files got zero requests, and Google says directly that you do not need AI text files to appear in its AI features. It has a real use in developer tooling, but for getting cited your time is better spent confirming AI crawlers can reach your pages.

### How long does LLM SEO take to show results?

It depends on the path. On the retrieval path, a reachable, well-structured page can be cited in live answers within days. On the training path, becoming part of a model's baseline knowledge takes model generations, often many months, and depends on broad third-party mentions you cannot rush.

### Can I stop AI from recommending my competitor instead of me?

Not directly. You cannot edit a model, and you cannot remove a competitor from its answers. What you can do is strengthen the signals the model reads about you, which means earning more third-party mentions, reviews, and citable content so the balance of evidence shifts in your favor over time.

### What do I do if an AI describes my business incorrectly?

Fix the sources it is reading, since you cannot edit the model. Correct outdated facts on your own pages, then work on the third-party pages, directories, and profiles the model pulls from, because a confident wrong answer often traces back to stale or conflicting information out on the web, though it can also be a plain model fabrication. Our guides on [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations) and [tracking AI brand sentiment](https://geotoolbox.ai/blog/ai-brand-sentiment) cover how to find and unwind these.

## Where to Start

Underneath the acronyms, LLM SEO is mostly disciplined SEO, plus digital PR, plus one thing the old playbook never had to worry about: whether the machines assembling the answer can reach you at all. The teams that win are not the ones buying a separate "GEO budget." They are the ones doing the fundamentals well and then covering the new ground, reachability and off-site presence, with the same rigor.

Start with the gate, because it is the cheapest thing to check and a common reason good work goes nowhere. Run your site through our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) to see which AI bots your robots.txt allows, then check your server logs and CDN for the ones that slip past it. If a retrieval crawler is blocked, fix that first. Everything else in this guide only pays off once the answer engines can see you.

## Sources

- GEO: Generative Engine Optimization - Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, 2023-2024 - `arxiv.org/abs/2311.09735`
- AI features and your website (best practices for SEO) - Google Search Central, updated 2026-07-10 - `developers.google.com/search/docs/fundamentals/ai-optimization-guide`
- Overview of OpenAI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User) - OpenAI - `developers.openai.com/api/docs/bots`
- Perplexity crawlers (PerplexityBot, Perplexity-User) - Perplexity - `docs.perplexity.ai/docs/resources/perplexity-crawlers`
- Does Anthropic crawl the web, and how to block it (ClaudeBot, Claude-User, Claude-SearchBot) - Anthropic - `support.claude.com/en/articles/8896518`
- Cloudflare just changed how AI crawlers scrape the internet (blocking AI crawlers by default) - Cloudflare, 2025-07-01 - `cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large`
- Your site, your rules: new AI traffic options (2026 category controls) - Cloudflare, 2026-07-01 - `blog.cloudflare.com/content-independence-day-ai-options`
- Declaring your AIndependence: block AI bots with one click - Cloudflare, 2024-07-03 - `blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click`
- Do LLMs.txt files actually work? (137,210-domain study, May 2026 data) - Ahrefs, 2026-06-15 - `ahrefs.com/blog/llmstxt-study`
- Sundar Pichai remarks, Google I/O (AI Overviews 2.5B monthly users, AI Mode 1B) - Google, 2026-05 - `blog.google/innovation-and-ai/sundar-pichai-io-2026`
